Agregación de backtests de estudios de caso con diagnósticos de selección
Resumen
Este cuaderno consulta registros de varios estudios de caso de trading y genera tablas entre conjuntos de datos para comparar señales, asignación, costes y riesgo. Informa de diagnósticos del grupo de configuraciones ordenadas por validación, como la diferencia de Sharpe entre la primera y la décima posición, el comportamiento por pliegue y otras medidas que muestran la estabilidad aparente de una selección. El proceso fija una especificación completa de estrategia, que abarca la señal, la asignación y el ajuste de riesgo, y empareja los resultados de validación con los de una muestra reservada reentrenada cuando están disponibles. Las restricciones sobre etiquetas y regímenes de ejecución mantienen determinadas comparaciones alineadas con el análisis previsto del estudio de caso.
Los artefactos exportados sirven para análisis posteriores de características, calidad de señales, asignación, costes y riesgo. El cuaderno también exige registros completos para una agregación total, ya que, de otro modo, las entradas parciales podrían confundirse con una comparación completa entre estudios. Las cifras derivadas de los registros son mediciones sujetas a incertidumbre, no clasificaciones definitivas: la selección de configuraciones puede ser frágil y no se deben reutilizar los resultados de la muestra reservada para elegir una estrategia. La cobertura depende de qué estudios de caso tengan backtests utilizables registrados, y las restricciones documentadas de etiquetas y regímenes determinan qué observaciones entran en las comparaciones seleccionadas.
Ideas clave
- La agregación de resultados de registros permite comparaciones entre conjuntos de datos, clases de activos y frecuencias de trading.
- La diferencia entre las configuraciones de validación mejor y peor clasificadas puede indicar la estabilidad de la selección.
- Empareja los resultados de validación y de la muestra reservada con la misma especificación completa de estrategia, incluidos la asignación y los ajustes de riesgo.
- Se necesita una cobertura completa de los registros para presentar los resultados como una síntesis total de varios estudios.
- Las métricas de los registros y los diagnósticos de selección aportan evidencia sujeta a incertidumbre, no clasificaciones universales de métodos.
Etiquetas
Texto completo
# Aggregate Synthesis
# Aggregate Synthesis
**Docker image**: `ml4t`
This notebook queries all 9 case study registries via `BacktestExplorer`
and builds cross-dataset comparison DataFrames for the remaining Ch20 notebooks.
**Data source**: `registry.db` per case study (no JSON files needed).
**Learning Objectives**:
- Query per-case-study backtest registries for signal, allocation, cost, and risk metrics
- Build cross-dataset comparison tables
- Export summary DataFrames for downstream notebooks (02–06)
**Book Reference**: Chapter 20, Section 20.1 (First-pass results across nine case studies)
**Prerequisites**: Case studies must have run Ch16–19 backtests.
```python
"""Ch20 Aggregate Synthesis — query registries and compare all 9 case studies."""
import json
import sqlite3
from functools import cache
from pathlib import Path
import polars as pl
import yaml
from IPython.display import Markdown, display
from case_studies.utils.backtest_explorer import BacktestExplorer
from case_studies.utils.benchmark import load_benchmark_returns
from case_studies.utils.strategy_analysis import (
allocation_method_of,
compute_cost_bps,
rank_one,
training_run_fitted_for_the_holdout,
)
from utils.paths import REPO_ROOT, get_case_study_dir, get_chapter_dir
```
```python
MAX_SYMBOLS = 0
# When non-empty, restricts the cross-CS iteration to the given subset.
# Used by the per-CS pipeline driver to populate `backtest_paired_metrics`
# for a single CS after its holdout has landed, without re-running the
# full 9-CS aggregation.
CASE_STUDIES: list[str] = []
# Test-only: in an isolated test registry, nasdaq's out-of-band cost-feasible
# selected configuration is absent, so its spine cannot resolve. Production leaves this False
# (a missing selection fails loudly); the test harness sets it True so cost/risk
# for such a case study are reported not-applicable instead of raising.
ALLOW_MISSING_SPINE = False
# A full-mode run overwrites the nine-case-study artifacts every downstream notebook and the
# chapter figures read. Refuse to start one unless all nine registries are present and carry
# backtests, so a partial aggregation cannot be published as a complete one. The test harness
# sets this False because its isolated registry is not the production store; a subset run
# (`CASE_STUDIES` non-empty) is the per-case-study driver path and is never full mode.
REQUIRE_ALL_REGISTRIES = True
```
```python
OUTPUT_DIR = get_chapter_dir(20) / "output"
OUTPUT_DIR.mkdir(exist_ok=True)
ALL_CASE_STUDIES = [
"etfs",
"crypto_perps_funding",
"nasdaq100_microstructure",
"sp500_equity_option_analytics",
"us_firm_characteristics",
# FX rank-1 is linear/ridge_a100.0 on fwd_ret_21d (val Sharpe +0.048,
# holdout +0.194), resolved after the 2026-06-01 DL-lookback fix. The
# earlier deep_learning/tcn selection (val +0.108 / holdout -1.59) was an
# artifact of gappy validation folds (lookback=60 warmup consumed each
# fold's head); those sets were purged and the clean lineage re-resolved.
# See backtest_audit.md and project_registry_hash_collisions.
"fx_pairs",
"cme_futures",
"sp500_options",
"us_equities_panel",
]
if CASE_STUDIES:
ALL_CASE_STUDIES = [cs for cs in ALL_CASE_STUDIES if cs in set(CASE_STUDIES)]
DISPLAY_NAMES = {
"etfs": "ETFs",
"crypto_perps_funding": "Crypto",
"nasdaq100_microstructure": "NASDAQ-100",
"sp500_equity_option_analytics": "S&P 500 Eq+Opt",
"us_firm_characteristics": "US Firms",
"fx_pairs": "FX Pairs",
"cme_futures": "CME Futures",
"sp500_options": "S&P 500 Options",
"us_equities_panel": "US Equities",
}
```
```python
ASSET_CLASS_MAP = {
"etfs": "equity_etf",
"crypto_perps_funding": "crypto",
"nasdaq100_microstructure": "equity_micro",
"sp500_equity_option_analytics": "equity_options",
"us_firm_characteristics": "equity_firm",
"fx_pairs": "fx",
"cme_futures": "futures",
"sp500_options": "options",
"us_equities_panel": "equity_panel",
}
FREQ_MAP = {
"etfs": "daily",
"crypto_perps_funding": "8h",
"nasdaq100_microstructure": "15min",
"sp500_equity_option_analytics": "daily",
"us_firm_characteristics": "monthly",
"fx_pairs": "daily",
"cme_futures": "daily",
"sp500_options": "daily",
"us_equities_panel": "daily",
}
```
## Load Registries
Create a `BacktestExplorer` for each case study that has a registry.
```python
explorers: dict[str, BacktestExplorer] = {}
configs: dict[str, dict] = {}
unreadable_registries: list[str] = []
empty_registries: list[str] = []
for cs in ALL_CASE_STUDIES:
try:
explorers[cs] = BacktestExplorer(cs)
setup_path = get_case_study_dir(cs) / "config" / "setup.yaml"
if setup_path.exists():
configs[cs] = yaml.safe_load(setup_path.read_text())
else:
configs[cs] = {}
summary = explorers[cs].summary()
total = sum(summary.values())
if total == 0:
empty_registries.append(cs)
print(f" [EMPTY] {cs}: registry.db present, zero backtest runs")
else:
print(f" [OK] {cs}: {total} backtests ({summary})")
except FileNotFoundError:
unreadable_registries.append(cs)
print(f" [MISSING] {cs}: no registry.db")
print(f"\nLoaded: {len(explorers)}/{len(ALL_CASE_STUDIES)} case studies")
```
### The full-mode precondition
This notebook overwrites the nine-case-study artifacts that notebooks 02 through 08 and the
chapter figures read. A run that finds only some of the nine registries produces an output
indistinguishable from a complete one, and on 2026-08-28 that is exactly what happened: a
synthesis run from a worktree carrying three registries stamped itself production and
published a holdout Sharpe under prose describing nine case studies. The push gate caught it;
the notebook did not.
So a full-mode run refuses to continue unless every case study named above has a readable
registry holding backtests. A subset run - `CASE_STUDIES` non-empty, the per-case-study
driver path that repopulates one case study's paired metrics after its holdout lands - is not
full mode and is not covered by the check.
```python
def refuse_partial_full_mode(
*,
expected: list[str],
subset: list[str],
unreadable: list[str],
empty: list[str],
enforce: bool = True,
) -> None:
"""Raise unless a full-mode run can see every registry it claims to aggregate.
A subset run - ``subset`` non-empty - is the per-case-study driver path and is not
full mode, so it is never refused. ``enforce`` is the seeded-test-registry escape and
is False in exactly one place, ``tests/overrides.yaml``.
"""
if not enforce or subset or not (unreadable or empty):
return
raise RuntimeError(
f"Full-mode synthesis needs all {len(expected)} registries present and holding "
"backtests, and this checkout does not have them. Refusing before anything is "
"written, because the artifacts this notebook overwrites are read as a complete "
"set covering every case study.\n"
f" no registry.db: {unreadable or 'none'}\n"
f" zero backtest runs: {empty or 'none'}\n"
"Run this once every case study has registered its backtests and the fleet has "
"stopped writing to them. Passing CASE_STUDIES is not a way round this: a subset "
"run repopulates one case study's paired metrics in its own registry and writes "
"none of the chapter-wide artifacts."
)
refuse_partial_full_mode(
expected=ALL_CASE_STUDIES,
subset=CASE_STUDIES,
unreadable=unreadable_registries,
empty=empty_registries,
enforce=REQUIRE_ALL_REGISTRIES,
)
```
## Top-Cluster Diagnostics
Rather than pre-committing to a single selected configuration per case study, we inspect the
*cluster* of top configurations on the validation split. A signal with genuine predictive
structure shows a thick top of the distribution: many configurations sit within a
fold-standard-error of the top-ranked Sharpe, and the implied pick is insensitive to small
perturbations in the selection rule. A thin cluster - a large gap between the top-ranked
configuration and the tenth - suggests the top result is closer to a tail draw than to a
stable optimum.
For each case study we report the top-ranked Sharpe, the tenth-ranked Sharpe where ten
configurations exist, the spread between them, the mean per-fold Sharpe, and the number of
folds in which the top-ranked configuration has positive Sharpe. These are measurements that
feed the downstream narrative.
### The selection rule
**A backtest's full strategy specification is the signal method, the allocation method and the
risk overlay taken together.** Naming all three is what makes a validation result and a
holdout result comparable, because it pins every stage rather than the signal alone.
Each case study's selected configuration is the highest-Sharpe validation backtest across
those three pipeline stages. The deployed holdout configuration is that same specification,
retrained on holdout data. When the holdout retrain produces no usable backtest at it -
degenerate predictions, a vol window that does not match the history available, a universe
filter that rejects the sample, or another generation failure - the rule falls back to the
next-highest validation Sharpe that does have a usable holdout, and so on until one succeeds.
The helper implementing that walk feeds the holdout query and the lineage resolver, which pin
each validation and holdout pair to one specification.
### Two restrictions, and why the selection needs both
**A label restriction**, so the cluster diagnostics and the Chapter 20 holdout retrain rank the
same thing. sp500_options trains a hold-to-maturity label with coherent option costs alongside
four fixed-horizon straddle labels priced through the vectorized path with a generic
basis-point cost. The Chapter 20 narrative uses the hold-to-maturity label as its
option-strategy reference, so restricting the cluster diagnostics to that same label keeps the
§20.1 top-cluster numbers aligned with the §20.5 and §20.6 narrative.
Those two section numbers are right, and they are recorded here because they will not look it.
Eleven references in this chapter pointed at sections that exist and do not carry what was
claimed, and a sweep for §20.5 in a chapter-20 notebook now finds this one and sees the same
shape. It is not the same: §20.5's Table 20.6 carries an sp500_options allocator row, and which
row that is depends on the label pinned here; §20.6 carries the option cost model, which is the
hold-to-maturity accounting rather than the basis-point sweep the other four labels get. Both
targets hold material this restriction decides. Do not retarget it.
**An execution-regime restriction**, because sp500_options is evaluated under the
O'Donovan-Yu (2025) cost-mitigation cascade, whose three rungs are a naive round trip, full
hold-to-maturity, and hold-to-maturity restricted to the liquid bottom-spread quintile. The
registered strategy is the third rung; the second is the demoted variant §18.8 discusses. The
first two rungs both carry the same universe filter, so filtering on that column
alone leaves `ORDER BY sharpe DESC LIMIT 1` free to pick whichever of the two happens to score
higher in the current data. Pinning the universe filter *and* the exit rule together is what
makes the selected row deterministic and coherent with hold-to-maturity. Case studies with no
entry here skip the filter altogether.
```python
# The rung pins are imported, not restated. This file used to carry its own copy of both
# predicates and of the dict around them, verbatim, and `paired_metrics.populate_paired_metrics`
# carries the other - both write `backtest_paired_metrics`, so a pin corrected on one side only
# would let one of them overwrite the other's rows with a differently-selected lineage. The
# duplication is how that drift happens, and the mirror keys beside each predicate
# (`universe_filter`, `exit_at_max_days`, `label`) exist for the SQL paths and `progression(...)`
# calls that cannot take a polars expression.
from case_studies.utils.paired_metrics import RUNG_PINS as _CLUSTER_RUNG_RESTRICTIONS # noqa: E402
from case_studies.utils.strategy_analysis import ( # noqa: E402
LABEL_RESTRICTIONS as _CLUSTER_LABEL_RESTRICTIONS,
)
from case_studies.utils.strategy_analysis import ( # noqa: E402
NoSelectableCandidates,
resolve_solvent_carrier,
selectable_validation_candidates,
)
@cache
def _canonical_carrier(cs: str) -> dict | None:
"""The configuration this case study reports, from the resolver that decides it.
This notebook built the same cross-stage rank-1 by hand in four places - concatenating
`explorer.best` over signal, allocation and risk_overlay, dropping benchmark families,
applying `LABEL_RESTRICTIONS` and `RUNG_PINS`, and taking the highest Sharpe. Three of
them wanted the winner and read it from here; the fourth, `_val_rank1_carrier`, walks the
whole field and takes it from `selectable_validation_candidates`, which is the same
ranking one step earlier. That is not the ranking the case studies report. `resolve_canonical_rank1_lineage` re-ranks the field
on exact common timestamp support whenever a conformal candidate is in it, because a
conformal allocator abstains until it is calibrated and books zeros over the abstention,
and it applies `UNIVERSE_RESTRICTIONS` and `CARRIER_PINS` besides. Measured 2026-09-18
against the nine canonical registries, the two rankings named different configurations on
two case studies: `fx_pairs` (`linear/ridge_a1000000.0` at Sharpe 0.4121
against `deep_learning/lstm_h64` at comparison Sharpe 0.3091) and
`nasdaq100_microstructure` (`deep_learning/nlinear` on fwd_ret_15m at 2.3001 against
`gbm/default_multiclass` on fwd_dir_15m at 2.4159). `spine_prediction_hash` is what
`05_portfolio_allocation` and Figure 20.7 pin their allocator comparison to, so on
`fx_pairs` the chapter compared allocators on a configuration the case study does not
report.
Returns None where the resolver finds nothing selectable, which is the state the
hand-built rankings reported as an empty frame. An insolvent or mis-calibrated carrier
still raises: it is a sweep to fix, not a case study to skip.
"""
try:
return resolve_solvent_carrier(cs)
except NoSelectableCandidates:
return None
def _retired(cs: str) -> frozenset[str]:
"""Identities a later generation retired, from the same helper `populate_paired_metrics`
uses. Both write `backtest_paired_metrics`, so a disagreement here would let a Chapter 20
run overwrite the corrected pairs with a retired lineage."""
from case_studies.utils.paired_metrics import _retired_prediction_hashes
return _retired_prediction_hashes(cs)
@cache
def _live_predictions(cs: str) -> list[str] | None:
"""What this case study currently publishes, or None when it declares no populations.
Membership, not the complement of retirement. A prediction no population ever listed has
not been retired by anyone, so ranking over "everything not retired" admits experimental
results the case study never published; ranking over the members in force does not.
Applied inside the query rather than to its result, because `best()` applies its SQL
`LIMIT top_n` first - a row filtered afterwards has already consumed a slot and can hide a
live candidate below the cut.
"""
from case_studies.research.population import published_members_at
published = published_members_at(get_case_study_dir(cs), member_kind="prediction")
if published is None:
return None
if not published:
# `best()` tests this argument for truthiness, so an empty list would read as "no
# filter" and rank everything. A study that declares populations and publishes
# nothing has nothing to report, which is a refusal rather than a wide-open ranking.
raise RuntimeError(f"{cs} declares populations but publishes no prediction identities")
return sorted(published)
def _best_live(explorer: "BacktestExplorer", cs: str, stage: str, top_n: int) -> pl.DataFrame:
"""`explorer.best` narrowed to what the case study still publishes."""
return explorer.best(stage=stage, top_n=top_n, prediction_hashes=_live_predictions(cs))
def _best_pinned(explorer: "BacktestExplorer", cs: str, stage: str, top_n: int) -> pl.DataFrame:
"""`explorer.best` for a stage, fetching enough rows that a rung-restricted
cohort survives the post-hoc predicate filter.
`best()` extracts `universe_filter` from `spec_json` in Python, after the
SQL `LIMIT top_n`. For nasdaq the pinned cost-feasible carrier sits below
the full-universe in-sample maxima, so a small `top_n` truncates it before
`_apply_rung_restriction` runs. Pull all rows for restricted case studies."""
live = _live_predictions(cs)
if cs in _CLUSTER_RUNG_RESTRICTIONS:
return explorer.best(stage=stage, top_n=1_000_000, prediction_hashes=live)
return explorer.best(stage=stage, top_n=top_n, prediction_hashes=live)
def _apply_rung_restriction(df: pl.DataFrame, cs: str) -> pl.DataFrame:
"""Filter `df` to the case study's pinned rung, if one is configured.
Returns the input untouched if no restriction applies. The helper
relies on `BacktestExplorer.best()` always emitting both
`universe_filter` and `exit_at_max_days` columns; if a future
schema regression drops them, the polars `filter` will raise a
column-not-found error rather than silently allowing the rank-1
selection to drift back to the cross-rung max."""
rung = _CLUSTER_RUNG_RESTRICTIONS.get(cs)
if rung is None or df.is_empty():
return df
return df.filter(rung["predicate"])
# There is no selected configuration pin here, and there is no mechanism for one.
# `_CARRIER_PIN_PREDICATES` held `"us_firm_characteristics": pl.col("config_name") ==
# "default_huber"` until 2026-08-25, copied from `case_studies.utils.strategy_analysis.CARRIER_PINS`
# and translated into a config-name predicate, under a "keep in sync" comment doing the job a
# mechanism should.
#
# It had not been in sync for a rebuild. Against the current registry `default_huber` is the
# WEAKEST of the ten configs that reached the allocation stage (48 validation backtests, best
# Sharpe 2.128, 2.075 average - tenth of ten), while the documented rule selects `leaves_63_mse`
# (59 backtests, 3.116). So this restricted one case study to its worst advanced configuration
# while every notebook inside that case study reported its best.
#
# Worse than the hash pin removed from `CARRIER_PINS` the same day, because a hash pin dies
# loudly: every hash changes when a sweep is rebuilt, so it resolves to nothing and stops. A
# config-name predicate survives the rebuild and keeps selecting, silently and wrongly.
#
# The mapping stayed empty behind an `_apply_carrier_pin` that could no longer fire, which is a
# second implementation of a rule nothing applied. A selected configuration restriction needed here
# again is `carrier_pins.carrier_config_name(cs)`, which resolves an owner's pin to its config
# through the registry - the thing the copy existed to avoid, and the thing that would have failed
# loudly rather than filtering to the wrong config.
def _progression_for(
explorer: "BacktestExplorer",
pred_hash: str,
cs: str,
) -> pl.DataFrame:
"""Call `progression()` with the case study's rung scope, if any."""
rung = _CLUSTER_RUNG_RESTRICTIONS.get(cs)
if rung is None:
return explorer.progression(pred_hash)
return explorer.progression(
pred_hash,
universe_filter=rung["universe_filter"],
exit_at_max_days=rung["exit_at_max_days"],
)
# Stages whose registry numbers should never be reported for a case study,
# either because the strategy makes the stage structurally meaningless (HTM
# short-straddle has no allocator choice, no bps cost sweep) or because the
# legacy registry contains deprecated entries that pre-date the strategy
# redesign. Consumed by both `build_backtest_rows` and the `synthesis_dict`
# sanitizer below so the in-notebook attrition funnel and the JSON artifact
# cannot drift on the same case study.
_STAGES_NOT_APPLICABLE: dict[str, set[str]] = {
"sp500_options": {"costs", "risk"},
"us_firm_characteristics": {"risk"},
"nasdaq100_microstructure": {"allocation", "costs", "risk"},
}
# Per-CS rationale strings for `not_applicable_reason` fields written into
# `synthesis_dict`. Keyed by (cs, stage) so two case studies that skip the
# same stage for different structural reasons render different prose.
_STAGE_NA_REASONS: dict[tuple[str, str], str] = {
("sp500_options", "allocation"): ("HTM short-straddle has fixed 1/n_roll cohort weighting"),
("sp500_options", "costs"): ("option costs use §18.8 bid-ask accounting, not bps sweep"),
("sp500_options", "risk"): ("HTM expiration structure sets risk profile"),
("us_firm_characteristics", "risk"): (
"vectorized-engine path; portfolio overlays purged 2026-05-17"
),
("nasdaq100_microstructure", "allocation"): (
"carrier is a signal-stage slot strategy; the slot mechanism is the sizing rule"
),
("nasdaq100_microstructure", "costs"): (
"timing-corrected broad carrier cost grid deferred to v3.1"
),
("nasdaq100_microstructure", "risk"): (
"timing-corrected broad carrier risk grid deferred to v3.1"
),
}
def _stage_applicable(cs: str, stage: str) -> bool:
"""Return False if `cs` has stage `stage` declared not applicable.
`stage` must be one of the canonical labels used by
``_STAGES_NOT_APPLICABLE`` itself (`allocation`, `costs`, `risk`).
Callers in this notebook always pass canonical literals."""
return stage not in _STAGES_NOT_APPLICABLE.get(cs, set())
```
```python
cluster_rows = []
for cs, explorer in explorers.items():
top = _best_pinned(explorer, cs, "signal", 200)
if top.is_empty() or top["sharpe"][0] is None:
continue
if "family" in top.columns:
top = top.filter(pl.col("family") != "benchmark")
label_restriction = _CLUSTER_LABEL_RESTRICTIONS.get(cs)
if label_restriction and "label" in top.columns:
top = top.filter(pl.col("label").is_in(list(label_restriction)))
top = _apply_rung_restriction(top, cs)
if top.is_empty() or top["sharpe"][0] is None:
continue
rank1 = top["sharpe"][0]
rank10 = top["sharpe"][9] if top.height >= 10 else None
spread = (rank1 - rank10) if rank10 is not None else None
# Fold-level stability for the rank-1 backtest
try:
bt_hash = top["backtest_hash"][0]
fold_df = explorer.fold_performance(bt_hash)
if not fold_df.is_empty():
mean_fold_sh = float(fold_df["sharpe"].mean())
se_fold_sh = (
float(fold_df["sharpe"].std(ddof=1) / (fold_df.height**0.5))
if fold_df.height > 1
else None
)
n_folds_pos = int(fold_df.filter(pl.col("sharpe") > 0).height)
n_folds = fold_df.height
else:
mean_fold_sh, se_fold_sh, n_folds_pos, n_folds = None, None, 0, 0
except Exception:
mean_fold_sh, se_fold_sh, n_folds_pos, n_folds = None, None, 0, 0
cluster_rows.append(
{
"case_study": DISPLAY_NAMES.get(cs, cs),
"cs_id": cs,
"rank1_sharpe": rank1,
"rank10_sharpe": rank10,
"rank1_rank10_spread": spread,
"mean_per_fold_sharpe": mean_fold_sh,
"fold_sharpe_se": se_fold_sh,
"n_folds_pos": n_folds_pos,
"n_folds": n_folds,
"n_configs": int(top.height),
}
)
cluster_df = pl.DataFrame(cluster_rows)
if not cluster_df.is_empty():
print("\n=== Rank-1 Cluster Diagnostics (validation) ===")
print(
cluster_df.select(
"case_study",
"rank1_sharpe",
"rank10_sharpe",
"rank1_rank10_spread",
"mean_per_fold_sharpe",
"fold_sharpe_se",
"n_folds_pos",
"n_folds",
"n_configs",
)
)
```
Read the table as a tuple: (rank1 Sharpe, rank10 Sharpe, spread, fold-SE,
folds-positive). A small rank1→rank10 spread relative to the fold-SE signals
a thick top-of-distribution. Folds-positive close to the total fold count
signals temporal stability. Both can be read off per case study without
collapsing the evidence into a single label.
## Overview Table
The nine case studies differ in asset class, rebalancing frequency, universe
size, and the cost assumption each one carries. The table records those four
properties so that a later result can be attributed to the setting it was
measured in rather than to the model that produced it.
```python
overview_rows = []
for cs, explorer in explorers.items():
setup = configs.get(cs, {})
cost_bps = compute_cost_bps(setup)
universe = setup.get("universe", {})
# A futures universe is sized in products rather than in assets, so
# cme_futures declares `n_products` where the others declare `n_assets`.
# Reading only the latter reported that case study as an empty universe.
n_assets = (
universe.get("n_assets")
or universe.get("n_products")
or len(universe.get("symbols", []))
or 0
)
primary_label = setup.get("labels", {}).get("primary", "")
families = explorer.compare_families(stage="signal")
overview_rows.append(
{
"case_study": DISPLAY_NAMES.get(cs, cs),
"cs_id": cs,
"asset_class": ASSET_CLASS_MAP.get(cs, "unknown"),
"frequency": FREQ_MAP.get(cs, "daily"),
"universe": n_assets,
"primary_label": primary_label,
"cost_bps": cost_bps,
"n_model_families": len(families) if not families.is_empty() else 0,
}
)
overview_df = pl.DataFrame(overview_rows)
overview_df.select("case_study", "asset_class", "frequency", "universe", "cost_bps")
```
The test bed covers equity ETFs, crypto perpetuals, intraday microstructure, equity plus
options, firm characteristics, FX, futures, pure options, and a broad equity panel. The
`cost_bps` column of the table above records the transaction-cost assumption each one
carries, so a later result can be read against the cost regime it was measured under.
## Model IC Comparison
Mean IC by model family across case studies, queried from `prediction_metrics`.
Each cell shows the average IC across all configurations within a family,
filtered to each case study's primary label.
```python
ic_rows = []
for cs, explorer in explorers.items():
case_dir = get_case_study_dir(cs)
db_path = case_dir / "run_log" / "registry.db"
if not db_path.exists():
continue
# Filter by primary label so IC values match book prose
primary_label = configs.get(cs, {}).get("labels", {}).get("primary", "")
db = sqlite3.connect(str(db_path))
# Best IC per family on primary label only
# Exclude causal_dml: it estimates treatment effects, not predictive IC.
# NOTE: best_ic and best_ic_daily are independent per-family MAXes — they may
# come from *different* predictions. This is intentional ("best daily IC in
# the family"), not "the daily IC of the best-by-fold model".
query = """
SELECT t.family, MAX(pm.ic_mean) AS best_ic,
MAX(pm.ic_mean_daily) AS best_ic_daily,
AVG(pm.ic_mean) AS mean_ic, COUNT(*) AS n_preds
FROM training_runs t
JOIN prediction_sets p ON t.training_hash = p.training_hash
JOIN prediction_metrics pm ON p.prediction_hash = pm.prediction_hash
WHERE p.split != 'holdout'
AND pm.ic_mean IS NOT NULL
AND t.family != 'causal_dml'
"""
params: tuple = ()
if primary_label:
query += " AND t.label = ?\n"
params = (primary_label,)
query += " GROUP BY t.family"
rows = db.execute(query, params).fetchall()
db.close()
for family, best_ic, best_ic_daily, mean_ic, n_preds in rows:
ic_rows.append(
{
"case_study": DISPLAY_NAMES.get(cs, cs),
"family": family,
"ic_mean": mean_ic,
"ic_best": best_ic,
"ic_best_daily": best_ic_daily,
"n_predictions": n_preds,
}
)
ic_df = pl.DataFrame(ic_rows)
```
```python
if not ic_df.is_empty():
ic_pivot = ic_df.pivot(on="family", index="case_study", values="ic_mean").sort("case_study")
else:
ic_pivot = pl.DataFrame()
```
### Mean IC by Model Family
```python
ic_pivot
```
Each cell is the mean of `ic_mean` over that family's non-holdout prediction sets, taken at
the primary label the case study declares in `setup.yaml`, with `causal_dml` excluded because
its runs are not fit to predict. Families are comparable within a row, since the label is
fixed across the row; they are not comparable across rows, because each case study declares a
different primary label, from a fifteen-minute forward return to a twenty-one-day one, and
prices a different instrument. A negative mean IC marks a case study where prediction is
difficult under that label rather than a defect in the family. §20.3 carries this table as
Table 20.4, and §20.6 works through how an option case study's raw IC translates into Sharpe
once single-name execution costs are charged.
## Backtest Comparison
Cross-dataset comparison of pipeline outcomes. Each row takes the highest-Sharpe result at
each stage **independently**, so the signal that tops one column may come from a different
model than the allocation that tops the next.
```python
def build_backtest_rows():
"""Build backtest comparison rows from all case study explorers."""
bt_rows = []
for cs, explorer in explorers.items():
summary = explorer.summary()
# Best signal-stage result — exclude benchmark families (equal_weight,
# etc.) since §20.4's model comparison is about trained models, not
# passive baselines. Also apply case-study label and universe-filter
# restrictions so the Ch20 rank-1 is HTM-coherent for sp500_options
# and pinned to the Rung-3 liquid subset, which is what
# `strategy_analysis.UNIVERSE_RESTRICTIONS` holds ({"sp500_options":
# "liquid"}). The Rung-2 full universe is retained only for the §18.8
# cascade comparison and never anchors the deployed carrier.
label_restriction = _CLUSTER_LABEL_RESTRICTIONS.get(cs)
signal_candidates = _best_pinned(explorer, cs, "signal", 200)
if not signal_candidates.is_empty() and "family" in signal_candidates.columns:
signal_candidates = signal_candidates.filter(pl.col("family") != "benchmark")
if (
label_restriction
and "label" in signal_candidates.columns
and not signal_candidates.is_empty()
):
signal_candidates = signal_candidates.filter(
pl.col("label").is_in(list(label_restriction))
)
signal_candidates = _apply_rung_restriction(signal_candidates, cs)
best_signal = signal_candidates.head(1)
signal_sharpe = best_signal["sharpe"][0] if not best_signal.is_empty() else None
best_source = best_signal["source"][0] if not best_signal.is_empty() else ""
# Carrier-pred pin for cost/risk. Case studies with a rung restriction
# (nasdaq cost-feasible ensemble) carry their headline cost/risk on the
# selected prediction only; the full-universe sweep rows are the
# Ch18/Ch19 cost-defeat demonstration and must not pool into the
# cross-case comparison. Other case studies pass None (no pin) and keep
# the registry-wide aggregation unchanged.
carrier_pred = (
best_signal["prediction_hash"][0]
if cs in _CLUSTER_RUNG_RESTRICTIONS and not best_signal.is_empty()
else None
)
# Best allocation-stage result (same filters). For case studies that
# declare the allocation stage not applicable (e.g. sp500_options HTM),
# the registry numbers come from deprecated runs, so report None to
# match the synthesis_dict sanitizer below.
if _stage_applicable(cs, "allocation"):
alloc_candidates = _best_live(explorer, cs, "allocation", 200)
if not alloc_candidates.is_empty() and "family" in alloc_candidates.columns:
alloc_candidates = alloc_candidates.filter(pl.col("family") != "benchmark")
if (
label_restriction
and "label" in alloc_candidates.columns
and not alloc_candidates.is_empty()
):
alloc_candidates = alloc_candidates.filter(
pl.col("label").is_in(list(label_restriction))
)
alloc_candidates = _apply_rung_restriction(alloc_candidates, cs)
best_alloc = alloc_candidates.head(1)
alloc_sharpe = best_alloc["sharpe"][0] if not best_alloc.is_empty() else None
# The allocator that produced `alloc_sharpe`, read from that row's own spec, so
# the name and the number describe one configuration. See
# `strategy_analysis.allocation_method_of`.
best_allocator = allocation_method_of(
cs, best_alloc["backtest_hash"][0] if not best_alloc.is_empty() else None
)
else:
alloc_sharpe = None
best_allocator = ""
# Cost sensitivity (gated by stage policy)
survives_costs = None
if _stage_applicable(cs, "costs"):
cost_df = explorer.cost_sensitivity(prediction_hash=carrier_pred)
if not cost_df.is_empty():
zero_cost = cost_df.filter(pl.col("cost_bps") == 0)
survives_costs = not zero_cost.is_empty() and zero_cost["sharpe"].max() > 0
# Risk overlay (gated by stage policy)
best_overlay = ""
managed_sharpe = None
if _stage_applicable(cs, "risk"):
risk_df = explorer.risk_impact(prediction_hash=carrier_pred)
if not risk_df.is_empty():
# rank_one, not a one-key sort: overlays that never trigger book the
# baseline Sharpe exactly, so ties at the top are ordinary here and a
# one-key sort would let frame order decide the name reported below.
best_risk_row = rank_one(risk_df, by="sharpe", name="risk_name")
best_overlay = best_risk_row["risk_name"][0]
managed_sharpe = best_risk_row["sharpe"][0]
# Spine rank-1 prediction_hash - the configuration the case study reports, taken
# from `_canonical_carrier` rather than ranked a second time here. Figure 20.7 and
# `05_portfolio_allocation` both read this value, and Ch20 prose Tables 20.5-20.7
# quote the resolver, so the two have to be one answer.
_carrier = _canonical_carrier(cs)
spine_pred_hash = _carrier["val_prediction_hash"] if _carrier else None
bt_rows.append(
{
"case_study": DISPLAY_NAMES.get(cs, cs),
"case_study_id": cs,
"spine_prediction_hash": spine_pred_hash,
"n_signal": summary.get("signal", 0),
"n_allocation": summary.get("allocation", 0),
"n_cost": summary.get("cost_sensitivity", 0),
"n_risk": summary.get("risk_overlay", 0),
"best_source": best_source,
"signal_sharpe": signal_sharpe,
"best_allocator": best_allocator,
"alloc_sharpe": alloc_sharpe,
"survives_costs": survives_costs,
"best_overlay": best_overlay,
"managed_sharpe": managed_sharpe,
}
)
return bt_rows
```
```python
bt_rows = build_backtest_rows()
```
```python
bt_df = pl.DataFrame(bt_rows)
print("\nPipeline Comparison:")
print(
bt_df.select(
"case_study",
"signal_sharpe",
"alloc_sharpe",
"survives_costs",
"managed_sharpe",
)
)
```
Read the baseline column of the table above for how many case studies enter the pipeline with
a positive baseline-stage Sharpe and which do not. Those counts move whenever a registry is
rebuilt, which is why they are in the table rather than in this sentence.
The lineage table below traces each selected prediction across the stages in the order the
backtests run: baseline, allocation, risk overlay, then the cost sweep charged against
whatever survived. A Sharpe that rises from one column to the next is what that stage added,
and only where the later stage carries the earlier one's configuration - the paired rows above
say which transitions meet that test. NASDAQ-100 is excluded from that comparison in v3.0
because its timing-corrected broad cost and risk grids are deferred to v3.1.
## Paired-Bootstrap Comparison vs Equal-Weight Benchmark
Each case study's selected baseline-stage backtest, under the same label, universe-filter and
rung restrictions used for the cluster diagnostics, is compared to its equal-weight benchmark
using a **paired stationary block bootstrap on daily strategy returns**. Block length is derived from
``setup.yaml.labels.{label}.rebalance_step`` (falling back to the optimal
block size, never below the label horizon). Reported quantities:
- ``sharpe_diff`` with a bootstrap confidence interval
- ``ret_diff``, the annualized return difference, with its confidence interval
- ``info_ratio`` of the daily-return difference
- ``prob_challenger_wins`` — bootstrap fraction in which challenger Sharpe
exceeds the benchmark
- ``p_value`` — two-sided bootstrap p-value for ``sharpe_diff = 0``
Results land in ``backtest_paired_metrics`` (per case study) and roll up
into the cross-dataset table below. Intervals are at the conventional confidence level the
bootstrap call sets. This is the right unit of uncertainty for the headline claim about the
selected configuration: the Sharpe **difference against the passive baseline that experienced
the same market conditions**, rather than the Sharpe alone.
```python
def _benchmark_returns_from_artifact(
cs: str, label: str, period: str = "overall"
) -> tuple[str, pl.DataFrame, str] | None:
"""Resolve the side-artifact equal-weight benchmark for ``(cs, label)``.
The benchmark is the daily-MTM EW reference series persisted by
``scripts/compute_vectorized_ew_benchmark.py`` at
``case_studies/{cs}/benchmark/{label}.parquet``. Single, well-defined
methodology per (cs, label) — no universe/rung/cadence ambiguity that
a registry-side ``family='benchmark'`` lookup would have to disambiguate.
``period`` selects the time window slice (``"overall"`` or ``"holdout"``)
applied by ``load_benchmark_returns``. Classification-label fallback to
the matching ``fwd_ret_*`` artifact applies in both periods.
Returns ``(synthetic_hash, returns_df, resolved_label)`` where
``synthetic_hash`` is a deterministic identifier safe to use as the PK
column in ``backtest_paired_metrics`` (which has no FK on
``benchmark_hash``) and ``resolved_label`` is the actual label whose
artifact was loaded — equal to ``label`` unless the classification
fallback fired, in which case it's the matching ``fwd_ret_*`` label.
Returns ``None`` if the artifact is missing.
"""
df = load_benchmark_returns(cs, label, period=period)
bench_label = label
if df.is_empty() or "ew_return" not in df.columns:
# Fallback: classification labels (``fwd_class_*``, ``fwd_dir_*``,
# ``fwd_tb_*``, ``fwd_carry_*``) share the same forecast window as
# their continuous counterpart (``fwd_ret_*``). The EW universe over
# the same window is identical regardless of the label being
# predicted, so map e.g. ``fwd_class_1m`` to ``fwd_ret_1m``.
fallback = None
for prefix in ("fwd_class_", "fwd_dir_", "fwd_tb_", "fwd_carry_"):
if label.startswith(prefix):
fallback = "fwd_ret_" + label[len(prefix) :]
break
if fallback is None:
return None
df = load_benchmark_returns(cs, fallback, period=period)
if df.is_empty() or "ew_return" not in df.columns:
return None
bench_label = fallback
suffix = "" if period == "overall" else f":{period}"
bench_hash = f"side_ew:{cs}:{bench_label}{suffix}"
return (
bench_hash,
df.select(
pl.col("timestamp").cast(pl.Date).alias("timestamp"),
pl.col("ew_return").cast(pl.Float64).alias("ret"),
),
bench_label,
)
def _aligned_returns(cs: str, h: str) -> pl.DataFrame | None:
"""Load and normalize a backtest's daily returns; columns ``[timestamp, ret]``."""
parquet = get_case_study_dir(cs) / "run_log" / "backtest" / h / "daily_returns.parquet"
if not parquet.exists():
return None
df = pl.read_parquet(parquet)
ret_col = next(
(c for c in ("daily_return", "ret", "return", "value") if c in df.columns),
df.columns[-1],
)
ts_col = next(
(c for c in ("timestamp", "date", "datetime") if c in df.columns),
df.columns[0],
)
return df.select(
pl.col(ts_col).cast(pl.Date).alias("timestamp"),
pl.col(ret_col).cast(pl.Float64).alias("ret"),
)
```
```python
import numpy as np
from case_studies.utils.uncertainty import (
SIGNAL_BASELINE_BY_CASE_STUDY,
STAGE_SEQUENCE,
compute_independent_diff_uncertainty,
compute_paired_uncertainty,
descends_from,
joint_returns,
)
def _min_paired_n(ppy: int) -> int:
"""Minimum series length for paired-bootstrap stability, frequency-aware.
The ~21 floor was written for daily cadences (about a month of obs).
Monthly case studies (e.g. ``us_firm_characteristics``) have ~12 holdout
observations by design, and ``compute_paired_uncertainty`` runs cleanly
on n=12. Scale the floor with ``ppy`` so monthly/weekly CSs aren't
blocked by a daily-tuned guard.
"""
if ppy <= 12: # monthly
return 6
if ppy <= 52: # weekly
return 12
return 21 # daily / 8h / intraday
# Distinguish skipped CSs from real failures so empty cross-dataset rollups
# aren't indistinguishable from a silent crash.
paired_rows: list[dict] = []
paired_skips: list[dict] = []
for cs, explorer in explorers.items():
# The leader is the configuration the case study reports, not a ranking rebuilt here.
# `_canonical_carrier` documents why the two are not the same ordering.
carrier = _canonical_carrier(cs)
if carrier is None:
paired_skips.append({"case_study": cs, "reason": "no_selectable_candidates"})
continue
leader_hash = carrier["val_backtest_hash"]
leader_label = carrier["label"]
if not leader_label:
paired_skips.append({"case_study": cs, "reason": "no_label_on_leader"})
continue
bench_resolution = _benchmark_returns_from_artifact(cs, leader_label)
if not bench_resolution:
paired_skips.append(
{"case_study": cs, "reason": f"no_benchmark_artifact_for_label:{leader_label}"}
)
continue
benchmark_hash, base, resolved_bench_label = bench_resolution
chal = _aligned_returns(cs, leader_hash)
if chal is None:
paired_skips.append({"case_study": cs, "reason": "no_challenger_returns_parquet"})
continue
ppy = {"daily": 252, "weekly": 52, "monthly": 12, "8h": 1095}.get(
FREQ_MAP.get(cs, "daily"), 252
)
min_n = _min_paired_n(ppy)
aligned = chal.join(base, on="timestamp", how="inner", suffix="_b")
if aligned.height < min_n:
paired_skips.append(
{"case_study": cs, "reason": f"insufficient_overlap:n={aligned.height}"}
)
continue
# A strategy against a benchmark: the leader's leading flat run is warmup before its
# first signal, not a position it held, so the sample starts where both are trading.
c_arr, b_arr = joint_returns(aligned["ret"].to_numpy(), aligned["ret_b"].to_numpy())
if c_arr.size < min_n:
paired_skips.append(
{"case_study": cs, "reason": f"insufficient_after_coerce:n={c_arr.size}"}
)
continue
paired = compute_paired_uncertainty(
c_arr,
b_arr,
periods_per_year=ppy,
case_study=cs,
label=leader_label,
n_boot=2000,
seed=42,
)
if not paired:
paired_skips.append({"case_study": cs, "reason": "paired_uncertainty_empty"})
continue
# Side-artifact benchmark — deterministic across (cs, label), no
# universe/rung ambiguity, no fallback-by-recency.
benchmark_kind = f"{SIGNAL_BASELINE_BY_CASE_STUDY.get(cs, 'equal_weight')}_side_artifact"
paired_rows.append(
{
"case_study": DISPLAY_NAMES.get(cs, cs),
"label": leader_label,
"benchmark_label": resolved_bench_label, # may differ from leader_label when classification fallback fired
"sharpe_diff": paired.get("sharpe_diff"),
"sharpe_diff_ci_lo": paired.get("sharpe_diff_ci95_lo"),
"sharpe_diff_ci_hi": paired.get("sharpe_diff_ci95_hi"),
"ret_diff": paired.get("ret_diff"),
"info_ratio": paired.get("info_ratio"),
"p_value": paired.get("p_value"),
"prob_wins": paired.get("prob_challenger_wins"),
"block": paired.get("bootstrap_block_length"),
"n_boot": paired.get("bootstrap_n"),
}
)
paired_df = pl.DataFrame(paired_rows)
if not paired_df.is_empty():
print("\n=== Paired Bootstrap: rank-1 vs equal-weight ===")
print(
paired_df.select(
"case_study",
"label",
"benchmark_label",
"sharpe_diff",
"sharpe_diff_ci_lo",
"sharpe_diff_ci_hi",
"info_ratio",
"prob_wins",
"p_value",
)
)
else:
print("\n=== Paired Bootstrap: rank-1 vs equal-weight ===")
print("No paired-bootstrap rows produced — see skip table below for reasons.")
if paired_skips:
print("\nSkipped case studies:")
for s in paired_skips:
print(f" - {s['case_study']:<32} {s['reason']}")
# Loud invariant — a 0/N or all-skipped outcome is now obvious in the
# notebook output instead of buried under "no paired-bootstrap rows."
print(f"\npaired={len(paired_rows)}/{len(explorers)}, skipped={len(paired_skips)}/{len(explorers)}")
```
Read each row as the selected challenger's annualized Sharpe **minus** the equal-weight
benchmark's, with a confidence interval from the paired stationary block bootstrap on the
daily-return difference; the information ratio summarizes the
excess-return-to-tracking-error ratio; ``prob_wins`` is the fraction of
bootstrap resamples in which the challenger beat the benchmark; ``p_value``
tests ``H0: sharpe_diff = 0``. A confident "the model adds skill over the
passive baseline" claim requires (i) the CI excludes zero, (ii) ``prob_wins``
close to 1, and (iii) a small ``p_value``. Cases where the CI straddles
zero are not failures — they signal that the apparent Sharpe gap is within
block-bootstrap sampling error and should be reported as such.
## Paired metrics — full coverage for strategy-analysis notebook
The block above populates the first pair type, the selected baseline signal against
equal-weight over the whole window. The strategy-analysis notebook (per-CS strategy notebooks)
requires five additional pair types per case study to render §2 (stage-
transition waterfall), §6 (holdout decay + holdout-vs-benchmark) and §7
(benchmark-aware diagnostics) without inline bootstrap recomputation.
The pair set:
1. selected signal (overall) ↔ equal-weight (overall) — populated above
2. selected signal (holdout) ↔ equal-weight (holdout window)
3. the selected configuration on holdout ↔ the same configuration on
validation (same lineage decay; min-length truncation since the
windows are disjoint)
4-6. one pair per consecutive stage transition the prediction actually has,
in ``STAGE_SEQUENCE`` order: allocation ↔ signal, risk-overlay ↔
allocation, cost-sensitivity ↔ risk-overlay. A case study that did not
run a stage yields fewer pairs, and a stage that does not carry the
previous stage's configuration yields none for that transition - the two
were selected independently and their difference is not a stage effect.
Pair #3 truncates both series to ``min(len(val), len(ho))`` to satisfy
``compute_paired_uncertainty``'s equal-length precondition. The CI is
interpreted as bootstrap resampling Sharpe in each window independently
and taking the difference; the truncation is preserved in the
``benchmark_kind`` value (``val_rank1_self`` always carries the truncation
caveat). All pairs use the same paired stationary block bootstrap helper.
```python
def _full_strategy_spec_from_backtest(db: sqlite3.Connection, bt_hash: str) -> dict | None:
"""Pull the full strategy spec dict (signal + allocation + risk) from
`bt_hash`'s spec_json. Returns None if the row is missing or signal has
no `method` field.
A backtest's full specification is the tuple (signal, allocation, risk). Pinning
the val→holdout pair on this full spec keeps the comparison apples-to-
apples; pinning on signal alone allows MAX(sharpe) to surface a holdout
row with a different allocation (e.g. conformal_weighted) or risk overlay
than the validation rank-1 carrier.
"""
row = db.execute(
"SELECT spec_json FROM backtest_runs WHERE backtest_hash = ?",
(bt_hash,),
).fetchone()
if not row:
return None
strat = json.loads(row[0]).get("strategy", {})
sig = strat.get("signal", {})
if not sig.get("method"):
return None
alloc = strat.get("allocation") or {}
risk = strat.get("risk") or {}
return {
"signal": {
"method": sig.get("method"),
"top_k": sig.get("top_k"),
"percentile": sig.get("percentile"),
},
"allocation": {
"method": alloc.get("method"),
"top_k": alloc.get("top_k"),
"long_short": alloc.get("long_short"),
},
"risk": {
"name": risk.get("name"),
},
}
def _val_rank1_carrier(cs: str) -> dict | None:
"""Return ``{'spec', 'prediction_hash'}`` for ``cs``'s validation rank-1 carrier.
The prediction hash is carried out alongside the spec because the holdout resolver
needs it: naming all three stages pins the configuration AND the checkpoint, and without
it a case study that registered several checkpoints against one strategy is ambiguous
and the resolver refuses. It was determinable all along - this walk had it in hand and
threw it away - so refusing there would have dropped a case study out of the
reader-facing holdout table for want of a value one line above.
The val rank-1 *full strategy* spec for ``cs`` — the
highest-Sharpe validation backtest across (signal, allocation,
risk_overlay) stages — walking candidates by val Sharpe descending until
one with a matching holdout backtest at the SAME full spec is found.
Implements the selection rule documented in §20.1: the deployed
holdout for each case study is the val rank-1 across all three pipeline
stages, retrained on holdout data; when retrain produces no usable
holdout at that full spec (degenerate predictions, vol-window mismatch,
universe filter rejection, etc.) the walk falls through to the next
candidate by val Sharpe. The first val candidate with a registered
holdout backtest at the same (signal, allocation, risk) tuple defines
the apples-to-apples carrier pair.
Returns None when no val candidate up to rank ~200 has a matching
holdout under the case study's label / rung restrictions.
"""
explorer = explorers.get(cs)
if explorer is None:
return None
# The field the resolver ranks, in the resolver's order, rather than a concat of
# `explorer.best` re-filtered here: the walk starts at the carrier `_canonical_carrier`
# names and falls through in the same order the resolver would. Every row is kept - no
# dedup by prediction_hash - because when the rank-1 configuration has no matching holdout
# retrain but a same-prediction lower-Sharpe variant (a different allocator or risk
# overlay) does, a dedup would jump to a different prediction instead of accepting the
# same-prediction variant as the apples-to-apples match.
try:
candidates = selectable_validation_candidates(cs)
except NoSelectableCandidates:
# The helper raises on an empty pool rather than returning one, and this walk's
# callers read `None` as "no holdout pair for this case study" - the state the
# hand-built ranking reported as an empty frame. A pool with nothing eligible in it
# is that state, not a reason to stop aggregating the other eight.
return None
label_restriction = _CLUSTER_LABEL_RESTRICTIONS.get(cs)
case_dir = get_case_study_dir(cs)
db_path = case_dir / "run_log" / "registry.db"
rung = _CLUSTER_RUNG_RESTRICTIONS.get(cs)
db = sqlite3.connect(str(db_path))
try:
for candidate in candidates[:200]:
bt_hash = candidate["backtest_hash"]
spec = _full_strategy_spec_from_backtest(db, bt_hash)
if spec is None:
continue
# The probe below asks whether THIS candidate has a holdout, so it matches the
# candidate's own configuration and checkpoint and not only its strategy spec.
#
# Matching the spec alone made the walk stop at a candidate whose own checkpoint
# had no holdout whenever a sibling checkpoint had one at the same spec. The
# resolver, handed that configuration, then finds nothing for it - and the walk has
# already stopped, so the case study reports no holdout while one exists for a
# later candidate. Advancing instead is what makes the fall-through the resolver
# no longer performs unnecessary rather than merely forbidden.
carrier_row = db.execute(
"""
SELECT t.family, t.config_name, t.label,
p.checkpoint_value, p.checkpoint_kind
FROM prediction_sets p
JOIN training_runs t ON t.training_hash = p.training_hash
WHERE p.prediction_hash = ?
""",
(candidate["prediction_hash"],),
).fetchone()
if carrier_row is None:
continue
spec_clauses, spec_params = _full_strategy_clauses(spec)
ho_clauses = ["p.split = 'holdout'"] + spec_clauses
if _retired(cs):
ho_clauses.append("p.prediction_hash NOT IN (SELECT value FROM json_each(?))")
ho_params: list[object] = list(spec_params)
if _retired(cs):
ho_params.append(json.dumps(sorted(_retired(cs))))
if label_restriction:
placeholders = ",".join("?" for _ in label_restriction)
ho_clauses.append(f"t.label IN ({placeholders})")
ho_params.extend(sorted(label_restriction))
if rung is not None:
ho_clauses.append(
"COALESCE(json_extract(b.spec_json, '$.strategy.signal.universe_filter'), 'full') = ?"
)
ho_params.append(rung["universe_filter"])
if rung["exit_at_max_days"] is None:
ho_clauses.append(
"json_extract(b.spec_json, '$.strategy.signal.exit_at_max_days') IS NULL"
)
else:
ho_clauses.append(
"json_extract(b.spec_json, '$.strategy.signal.exit_at_max_days') = ?"
)
ho_params.append(rung["exit_at_max_days"])
# `t.spec_json` rather than `1`, and no LIMIT: the probe has to apply the same
# eligibility test the resolver applies, and that test is not expressible in SQL.
#
# A model fitted on the validation folds can publish predictions over the holdout
# window, so `p.split = 'holdout'` with a non-null Sharpe is not enough to make a
# row a holdout result. The resolver drops those through
# `training_run_fitted_for_the_holdout`; a probe that admitted them would stop the
# walk at a candidate whose only holdout is validation-fitted, the resolver would
# then find nothing eligible for it, and the case study would report no holdout
# while a later candidate had a real one.
probe_rows = db.execute(
f"""
SELECT t.spec_json FROM prediction_sets p
JOIN training_runs t ON p.training_hash = t.training_hash
JOIN backtest_runs b ON p.prediction_hash = b.prediction_hash
AND b.stage IN ('signal','allocation','risk_overlay','holdout')
JOIN backtest_metrics bm ON b.backtest_hash = bm.backtest_hash
WHERE {" AND ".join(ho_clauses)}
AND t.family = ?
AND t.config_name = ?
AND t.label = ?
AND p.checkpoint_value IS ?
AND p.checkpoint_kind IS ?
AND bm.sharpe IS NOT NULL
""",
ho_params + list(carrier_row),
).fetchall()
row = any(training_run_fitted_for_the_holdout(probe[0]) for probe in probe_rows)
if row:
return {"spec": spec, "prediction_hash": candidate["prediction_hash"]}
finally:
db.close()
return None
def _full_strategy_clauses(spec: dict | None) -> tuple[list[str], list[object]]:
"""Build SQL WHERE clauses + params that pin a backtest row to the full
strategy spec (signal + allocation + risk). Empty list returned when
spec is None (no constraint).
Pinning on the full spec ensures `MAX(sharpe)` over candidate holdout
backtests cannot surface a different allocator (e.g. conformal_weighted
when val rank-1 was score_weighted) or a different risk overlay than
the validation carrier — the val→holdout pair stays apples-to-apples
on the full pipeline configuration, not just the signal.
"""
if not spec:
return [], []
clauses: list[str] = []
params: list[object] = []
sig = spec.get("signal") or {}
method = sig.get("method")
if method is None:
clauses.append("json_extract(b.spec_json, '$.strategy.signal.method') IS NULL")
else:
clauses.append("json_extract(b.spec_json, '$.strategy.signal.method') = ?")
params.append(method)
top_k = sig.get("top_k")
if top_k is None:
clauses.append("json_extract(b.spec_json, '$.strategy.signal.top_k') IS NULL")
else:
clauses.append("CAST(json_extract(b.spec_json, '$.strategy.signal.top_k') AS INTEGER) = ?")
params.append(int(top_k))
pct = sig.get("percentile")
if pct is None:
clauses.append("json_eSe muestra íntegramente con atribución según la licencia de la fuente. Licencia: MIT
Este resumen lo redactó el agente de investigación de Stratmill a partir del original; no es una copia de la fuente.