بیک ٹیسٹ انتخابی مراحل میں جوڑا وار بُوٹ اسٹریپ موازنے
خلاصہ
یہ دستاویز اسٹریٹیجی انتخاب کے مراحل میں جوڑا وار کارکردگی موازنوں کا ایک تیارکنندہ بیان کرتی ہے۔ یہ سگنل کے انتخابات اور مساوی وزن کے حوالوں، ہولڈ آؤٹ اور توثیقی انتخابات، نیز تخصیص، لاگت کی حساسیت اور خطرے کے اوورلے کی تبدیلیوں کا تقابل کرتی ہے۔ زیادہ تر موازنے جوڑا وار ساکن بلاک بُوٹ اسٹریپ استعمال کرتے ہیں، جو ہم آہنگ منافع کی سیریز کا تعلق برقرار رکھتا ہے۔ توثیق اور ہولڈ آؤٹ موازنے میں مشاہدات کے غیر متداخل ہونے کے باعث یہ الگ الگ کھڑکیوں سے نمونے لیتا ہے؛ اس وقفے سے ان کھڑکیوں کا فرق ناپا جاتا ہے، کسی ایک مستقل اسٹریٹیجی برتری کا ثبوت نہیں۔
یہ تیارکنندہ عملی انتخابی پابندیاں بھی سنبھالتا ہے، مثلاً مخصوص کیس کے لیبل یا اسٹریٹیجی درجے کی پابندیاں، مشاہدے کی فریکوئنسی کے مطابق کم از کم نمونہ حجم، گم شدہ منافع کے آرٹیفیکٹس، اور مکمل رَن کے بعد پرانی رجسٹری قطاروں کی صفائی۔ یہ تفصیلات مختلف کیس اسٹڈیز میں یکساں رپورٹنگ میں مدد دیتی ہیں، مگر ماڈیول خود غیر یقینی کی پیمائش کا بنیادی ڈھانچہ ہے، ٹریڈنگ اسٹریٹیجی یا اس تجرباتی دعوے کا ثبوت نہیں کہ کوئی ایک انتخابی طریقہ جیتتا ہے۔ بُوٹ اسٹریپ وقفے دستیاب منافع کی تاریخ اور موازنے کے ڈیزائن پر منحصر ہیں؛ الگ ادوار سے کارکردگی کے تخمینے شماریاتی طور پر آزاد نہیں ہو جاتے۔
اہم خیالات
- جوڑا وار ساکن بلاک بُوٹ اسٹریپ موازنہ کارکردگی کے فرق کی غیر یقینی جانچنے کے لیے ہم آہنگ منافع کے مشاہدات استعمال کرتا ہے۔
- الگ توثیقی اور ہولڈ آؤٹ کھڑکیوں کے لیے آزاد نمونے درکار ہیں، اور ان کا وقفہ دونوں کھڑکیوں کا فرق ناپتا ہے۔
- یہ موازنہ سگنل، تخصیص، لاگت اور رسک اوورلے کے انتخابی مراحل میں کارکردگی کی تبدیلیاں دیکھتے ہیں۔
- فریکوئنسی کے مطابق کم از کم نمونہ حدیں کیس اسٹڈیز میں مشاہدات کی مختلف تعداد کو ملحوظ رکھتی ہیں۔
- بُوٹ اسٹریپ نتائج مقررہ نمونے اور ڈیزائن کے تحت غیر یقینی بتاتے ہیں؛ وہ مستقل ٹریڈنگ برتری ثابت نہیں کرتے۔
ٹیگز
مکمل متن
# paired_metrics.py
```py
"""Compute-and-register the per-case-study ``backtest_paired_metrics`` table.
In-repo home of the paired-bootstrap producer that the strategy-analysis
notebook (``NN_strategy_analysis.py``) runs so a reader who never touches
Chapter 20 still lands a populated ``backtest_paired_metrics`` table. It renders
§2 (stage-transition waterfall), §6 (holdout decay + holdout-vs-benchmark) and
§7 (benchmark-aware diagnostics) without inline bootstrap recomputation.
The logic is extracted verbatim from the two producer loops in
``20_strategy_synthesis/01_aggregate_synthesis.py``; the only change is that the
per-case-study selection config (label restriction, rung pin, carrier pin,
frequency) is passed in as arguments rather than read from Ch20 module globals,
so a single case study can run standalone. The Chapter-20 aggregate can later
call this function per case study and become a pure reader.
Six pair types are produced per case study:
1. signal rank-1 (overall) ↔ equal-weight (overall)
2. cross-stage rank-1 holdout ↔ equal-weight (holdout window)
3. holdout rank-1 ↔ validation rank-1 of the same lineage (disjoint windows)
4. allocation rank-1 ↔ signal rank-1 (same window, stage transition)
5. cost-sensitivity rank-1 ↔ allocation rank-1 (same window)
6. risk-overlay rank-1 ↔ cost-sensitivity rank-1 (same window)
All pairs use the paired stationary block bootstrap
(``compute_paired_uncertainty``); pair #3 uses independent per-window draws
(``compute_independent_diff_uncertainty``) because its two windows share no
observations, so there is no difference series to pair on. Disjointness removes the
pairing; it does not make the two Sharpes independent, and the interval that comes
back is calibrated for the gap between those two windows rather than for the
strategy having one edge across both. See that function for the measurement.
"""
from __future__ import annotations
import json
import sqlite3
from functools import cache
from pathlib import Path
import numpy as np
import polars as pl
from case_studies.utils.analytics import DISPLAY_NAMES
from case_studies.utils.backtest_explorer import BacktestExplorer
from case_studies.utils.benchmark import load_benchmark_returns
from case_studies.utils.notebook_contracts import degenerate_prediction_hashes
from case_studies.utils.registry.registration import register_paired_metrics
from case_studies.utils.strategy_analysis import (
is_refit_of,
training_run_fitted_for_the_holdout,
)
from case_studies.utils.uncertainty import (
SIGNAL_BASELINE_BY_CASE_STUDY,
STAGE_SEQUENCE,
CarrierScope,
EntireRegistry,
NoCarrier,
PredictionScope,
compute_independent_diff_uncertainty,
compute_paired_uncertainty,
descends_from,
joint_returns,
)
from utils.paths import get_case_study_dir
# Cross-stage rank-1 pooling stages - mirrors strategy_analysis.SELECTION_STAGES.
_PAIRED_STAGES = ("signal", "allocation", "risk_overlay")
def _min_paired_n(ppy: int) -> int:
"""Minimum series length for paired-bootstrap stability, frequency-aware.
The ~21 floor was written for daily cadences (about a month of obs).
Monthly case studies (e.g. ``us_firm_characteristics``) have ~12 holdout
observations by design, and ``compute_paired_uncertainty`` runs cleanly
on n=12. Scale the floor with ``ppy`` so monthly/weekly CSs aren't
blocked by a daily-tuned guard.
"""
if ppy <= 12: # monthly
return 6
if ppy <= 52: # weekly
return 12
return 21 # daily / 8h / intraday
# The two case studies whose canonical strategy is pinned to one rung of a cascade, mirroring
# `20_strategy_synthesis/01_aggregate_synthesis.py::_CLUSTER_RUNG_RESTRICTIONS`. Held here as
# well because a case study's own strategy-analysis notebook has to make the same selection
# without importing a chapter, and `tests/test_rung_pins_match_chapter_20.py` fails if the two
# definitions drift.
#
# sp500_options: rung-1 (mid-to-mid bps) and rung-2 (full-universe HTM) both carry
# `universe_filter="full"`, so filtering on the universe alone leaves `ORDER BY sharpe DESC
# LIMIT 1` free to pick whichever rung is higher in current data. The pin combines the universe
# with `exit_at_max_days` so the rank-1 row is deterministic and HTM-coherent.
#
# nasdaq100_microstructure: the cost-feasible sweep on the primary label, matched on design
# attributes any registry can satisfy and chosen before the holdout was opened.
#
# The pin used to name `family == "ensemble"` as well. The mean-forecast ensemble existed
# because the per-model baseline on this case study was not worth reporting, and that is no
# longer the case: measured 2026-09-14 on the cost-feasible pool, `deep_learning/nlinear` on
# `fwd_ret_15m` reaches +2.300 against the ensemble's +0.566, so pinning to the ensemble
# anchored every paired comparison to the weakest thing in the pool. The family clause is
# gone; the ensemble rows stay in the registry and stay selectable, they are simply no
# longer the only thing the pin can choose.
#
# The label is part of the pin and was not always, and it does more work now that family is
# not. The pool spans four declared labels, so universe alone leaves `ORDER BY sharpe DESC
# LIMIT 1` free to choose among them - `fwd_dir_15m` reaches +2.416, above the primary
# label's best - and the rank-1 rung would move onto a label this case study is not featured
# on with nothing announcing it. A pin that omits a dimension selects along it silently.
#
# `fwd_ret_15m` is what the book prints for this case study, in Table 11.6
# (`NASDAQ-100 15m | fwd_ret_15m`), in Chapter 12's case-study table (`15 minutes | forward
# return`) and in Chapter 13's (`15 minutes | NLinear`). `config/setup.yaml::labels.primary`
# says the same, and `tests/test_rung_pin_label.py` asserts the two do not drift apart.
RUNG_PINS: dict[str, dict] = {
"sp500_options": {
"predicate": (pl.col("universe_filter") == "liquid") & pl.col("exit_at_max_days").is_null(),
"universe_filter": "liquid",
"exit_at_max_days": None,
},
"nasdaq100_microstructure": {
"predicate": (pl.col("universe_filter") == "cost_feasible")
& (pl.col("label") == "fwd_ret_15m"),
"universe_filter": "cost_feasible",
"exit_at_max_days": None,
# Mirrors the predicate for the SQL paths that cannot take a polars expression.
"label": "fwd_ret_15m",
},
}
def rung_for(cs: str) -> dict | None:
"""The rung this case study is pinned to, or None where it is not pinned."""
return RUNG_PINS.get(cs)
def _best_for_rung(
explorer: BacktestExplorer,
stage: str,
rung: dict | None,
top_n: int = 2000,
prediction_hashes: list[str] | None = None,
) -> pl.DataFrame:
"""``explorer.best`` for a stage, fetching enough rows that the pin survives.
``best()`` reads ``universe_filter`` out of ``spec_json`` in Python, and truncates to
``top_n`` after that; ``_apply_rung_restriction`` runs later still. For nasdaq the pinned
cost-feasible carrier sits below the full-universe in-sample maxima, so a small ``top_n``
truncates it before the predicate is ever applied and the pin silently selects nothing -
or, worse, the best surviving row that was never the carrier.
A pinned case study therefore asks for every row, which ``best`` now spells ``top_n=0``.
It was a literal million until ``sweep_config.top_n_cap`` gave 0 that meaning, and a cohort
past a million would have been truncated rather than refused.
Ch20 solves this with the same widening (`_best_pinned`); the extraction into this module
dropped it, which is why every pinned selection here has to go through this helper rather
than call ``explorer.best`` directly.
"""
return explorer.best(
stage=stage,
top_n=0 if rung is not None else top_n,
prediction_hashes=prediction_hashes,
)
def _apply_rung_restriction(df: pl.DataFrame, rung: dict | None) -> pl.DataFrame:
"""Filter ``df`` to the case study's pinned rung, if one is configured.
Returns the input untouched if ``rung`` is None. The helper relies on
``BacktestExplorer.best()`` always emitting both ``universe_filter`` and
``exit_at_max_days`` columns; if a future schema regression drops them, the
polars ``filter`` raises column-not-found rather than silently letting the
rank-1 selection drift back to the cross-rung max.
"""
if rung is None or df.is_empty():
return df
return df.filter(rung["predicate"])
def _benchmark_returns_from_artifact(
cs: str, label: str, period: str = "overall"
) -> tuple[str, pl.DataFrame, str] | None:
"""Resolve the side-artifact equal-weight benchmark for ``(cs, label)``.
The benchmark is the daily-MTM EW reference series persisted by
``scripts/compute_vectorized_ew_benchmark.py`` at
``case_studies/{cs}/benchmark/{label}.parquet``. Single, well-defined
methodology per (cs, label). ``period`` selects the window slice
(``"overall"`` or ``"holdout"``). Classification-label fallback to the
matching ``fwd_ret_*`` artifact applies in both periods.
Returns ``(synthetic_hash, returns_df, resolved_label)`` or ``None`` when
the artifact is missing. ``synthetic_hash`` is a deterministic identifier
safe as the PK column in ``backtest_paired_metrics`` (no FK on
``benchmark_hash``).
"""
df = load_benchmark_returns(cs, label, period=period)
bench_label = label
if df.is_empty() or "ew_return" not in df.columns:
fallback = None
for prefix in ("fwd_class_", "fwd_dir_", "fwd_tb_", "fwd_carry_"):
if label.startswith(prefix):
fallback = "fwd_ret_" + label[len(prefix) :]
break
if fallback is None:
return None
df = load_benchmark_returns(cs, fallback, period=period)
if df.is_empty() or "ew_return" not in df.columns:
return None
bench_label = fallback
suffix = "" if period == "overall" else f":{period}"
bench_hash = f"side_ew:{cs}:{bench_label}{suffix}"
return (
bench_hash,
df.select(
pl.col("timestamp").cast(pl.Date).alias("timestamp"),
pl.col("ew_return").cast(pl.Float64).alias("ret"),
),
bench_label,
)
def _aligned_returns(cs: str, h: str) -> pl.DataFrame | None:
"""Load and normalize a backtest's daily returns; columns ``[timestamp, ret]``."""
parquet = get_case_study_dir(cs) / "run_log" / "backtest" / h / "daily_returns.parquet"
if not parquet.exists():
return None
df = pl.read_parquet(parquet)
ret_col = next(
(c for c in ("daily_return", "ret", "return", "value") if c in df.columns),
df.columns[-1],
)
ts_col = next(
(c for c in ("timestamp", "date", "datetime") if c in df.columns),
df.columns[0],
)
return df.select(
pl.col(ts_col).cast(pl.Date).alias("timestamp"),
pl.col(ret_col).cast(pl.Float64).alias("ret"),
)
def _full_strategy_spec_from_backtest(db: sqlite3.Connection, bt_hash: str) -> dict | None:
"""Pull the full strategy spec dict (signal + allocation + risk) from
``bt_hash``'s spec_json. Returns None if the row is missing or signal has
no ``method`` field.
The carrier of a backtest is the tuple (signal, allocation, risk). Pinning
the val→holdout pair on this full spec keeps the comparison apples-to-
apples; pinning on signal alone allows MAX(sharpe) to surface a holdout
row with a different allocation or risk overlay than the validation rank-1.
"""
row = db.execute(
"SELECT spec_json FROM backtest_runs WHERE backtest_hash = ?",
(bt_hash,),
).fetchone()
if not row:
return None
strat = json.loads(row[0]).get("strategy", {})
sig = strat.get("signal", {})
if not sig.get("method"):
return None
alloc = strat.get("allocation") or {}
risk = strat.get("risk") or {}
return {
"signal": {
"method": sig.get("method"),
"top_k": sig.get("top_k"),
"percentile": sig.get("percentile"),
},
"allocation": {
"method": alloc.get("method"),
"top_k": alloc.get("top_k"),
"long_short": alloc.get("long_short"),
},
"risk": {
"name": risk.get("name"),
},
}
def _full_strategy_clauses(spec: dict | None) -> tuple[list[str], list[object]]:
"""Build SQL WHERE clauses + params pinning a backtest row to the full
strategy spec (signal + allocation + risk). Empty list when spec is None.
Pinning on the full spec ensures ``MAX(sharpe)`` over candidate holdout
backtests cannot surface a different allocator or risk overlay than the
validation carrier — the val→holdout pair stays apples-to-apples on the
full pipeline configuration, not just the signal.
"""
if not spec:
return [], []
clauses: list[str] = []
params: list[object] = []
sig = spec.get("signal") or {}
method = sig.get("method")
if method is None:
clauses.append("json_extract(b.spec_json, '$.strategy.signal.method') IS NULL")
else:
clauses.append("json_extract(b.spec_json, '$.strategy.signal.method') = ?")
params.append(method)
top_k = sig.get("top_k")
if top_k is None:
clauses.append("json_extract(b.spec_json, '$.strategy.signal.top_k') IS NULL")
else:
clauses.append("CAST(json_extract(b.spec_json, '$.strategy.signal.top_k') AS INTEGER) = ?")
params.append(int(top_k))
pct = sig.get("percentile")
if pct is None:
clauses.append("json_extract(b.spec_json, '$.strategy.signal.percentile') IS NULL")
else:
clauses.append(
"CAST(json_extract(b.spec_json, '$.strategy.signal.percentile') AS REAL) = ?"
)
params.append(float(pct))
alloc = spec.get("allocation") or {}
am = alloc.get("method")
if am is None:
clauses.append("json_extract(b.spec_json, '$.strategy.allocation.method') IS NULL")
else:
clauses.append("json_extract(b.spec_json, '$.strategy.allocation.method') = ?")
params.append(am)
ak = alloc.get("top_k")
if ak is None:
clauses.append("json_extract(b.spec_json, '$.strategy.allocation.top_k') IS NULL")
else:
clauses.append(
"CAST(json_extract(b.spec_json, '$.strategy.allocation.top_k') AS INTEGER) = ?"
)
params.append(int(ak))
als = alloc.get("long_short")
if als is None:
clauses.append("json_extract(b.spec_json, '$.strategy.allocation.long_short') IS NULL")
else:
clauses.append(
"CAST(json_extract(b.spec_json, '$.strategy.allocation.long_short') AS INTEGER) = ?"
)
params.append(int(bool(als)))
risk = spec.get("risk") or {}
risk_name = risk.get("name")
if risk_name is None:
clauses.append("json_extract(b.spec_json, '$.strategy.risk.name') IS NULL")
else:
clauses.append("json_extract(b.spec_json, '$.strategy.risk.name') = ?")
params.append(risk_name)
return clauses, params
@cache
def _retired_prediction_hashes(cs: str) -> frozenset[str]:
"""Every prediction identity a later generation retired, on either split.
Retirement is recorded member-wise on the *validation* population, and a prediction
identity includes its split, so a retired validation hash never equals its holdout
sibling and cannot filter it directly. What identifies the same model state across the
two is the training run together with the checkpoint it was scored at, so this expands
the recorded set along that key.
Member-wise, not run-wise: a population that moved one checkpoint of a run and kept
another has retired one checkpoint, and the sibling that did not move is still current.
``superseded_members_at`` rather than ``superseded_members``, because the latter takes a
``Study`` and every ``Study.open`` branch ends in ``activate()``, which clears
``ML4T_OUTPUT_DIR`` and would re-point a preview or isolated workspace at the released
registry - answering for a different registry than these queries read.
"""
from case_studies.research.population import superseded_members_at
case_dir = get_case_study_dir(cs)
db_path = case_dir / "run_log" / "registry.db"
if not db_path.exists():
return frozenset()
recorded = superseded_members_at(case_dir, member_kind="prediction")
if not recorded:
return frozenset()
db = sqlite3.connect(str(db_path))
try:
rows = db.execute(
"SELECT prediction_hash, training_hash, checkpoint_kind, checkpoint_value "
"FROM prediction_sets"
).fetchall()
finally:
db.close()
retired_states = {
(training_hash, checkpoint_kind, checkpoint_value)
for prediction_hash, training_hash, checkpoint_kind, checkpoint_value in rows
if prediction_hash in recorded
}
return frozenset(
prediction_hash
for prediction_hash, training_hash, checkpoint_kind, checkpoint_value in rows
if (training_hash, checkpoint_kind, checkpoint_value) in retired_states
)
def _val_rank1_carrier(
cs: str,
explorer: BacktestExplorer,
*,
label_restriction: frozenset[str] | None,
rung: dict | None,
prediction_hashes: list[str] | None = None,
retired_hashes: frozenset[str] | None = None,
) -> dict | None:
"""Return the val rank-1 *full strategy* spec for ``cs`` — the
highest-Sharpe validation backtest across (signal, allocation,
risk_overlay) stages — walking candidates by val Sharpe descending until
one with a matching holdout backtest at the SAME full spec is found.
Returns None when no val candidate up to rank ~200 has a matching holdout
under the case study's label / rung restrictions.
"""
cand = pl.concat(
[
_best_for_rung(explorer, s, rung, prediction_hashes=prediction_hashes)
for s in ("signal", "allocation", "risk_overlay")
],
how="diagonal_relaxed",
)
if cand.is_empty() or "backtest_hash" not in cand.columns:
return None
cand = _eligible_candidates(cs, cand, label_restriction=label_restriction, rung=rung)
if cand.is_empty():
return None
# Do NOT dedup by prediction_hash here — the walk needs every registered
# (signal, allocation, risk_overlay) tuple so a same-prediction lower-sharpe
# variant can serve as the apples-to-apples carrier.
cand = cand.sort("sharpe", descending=True)
case_dir = get_case_study_dir(cs)
db_path = case_dir / "run_log" / "registry.db"
db = sqlite3.connect(str(db_path))
try:
for i in range(min(cand.height, 200)):
bt_hash = cand["backtest_hash"][i]
spec = _full_strategy_spec_from_backtest(db, bt_hash)
if spec is None:
continue
spec_clauses, spec_params = _full_strategy_clauses(spec)
ho_clauses = ["p.split = 'holdout'"] + spec_clauses
ho_params: list[object] = list(spec_params)
# The same exclusion `_holdout_lineage_for` applies. Without it a retired holdout
# makes this candidate look eligible, the walk stops here, and that call then
# filters the row out and returns nothing - instead of advancing to the next live
# candidate that does have a holdout.
if retired_hashes:
ho_clauses.append("p.prediction_hash NOT IN (SELECT value FROM json_each(?))")
ho_params.append(json.dumps(sorted(retired_hashes)))
if label_restriction:
placeholders = ",".join("?" for _ in label_restriction)
ho_clauses.append(f"t.label IN ({placeholders})")
ho_params.extend(sorted(label_restriction))
if rung is not None:
ho_clauses.append(
"COALESCE(json_extract(b.spec_json, '$.strategy.signal.universe_filter'), 'full') = ?"
)
ho_params.append(rung["universe_filter"])
if rung["exit_at_max_days"] is None:
ho_clauses.append(
"json_extract(b.spec_json, '$.strategy.signal.exit_at_max_days') IS NULL"
)
else:
ho_clauses.append(
"json_extract(b.spec_json, '$.strategy.signal.exit_at_max_days') = ?"
)
ho_params.append(rung["exit_at_max_days"])
# The probe asks whether THIS candidate has a holdout, so it matches the
# candidate's own configuration and checkpoint, not only its strategy spec, and
# applies the same eligibility test `_holdout_lineage_for` applies.
#
# Spec alone stopped the walk wherever a SIBLING checkpoint had a holdout at the
# same spec. The caller then pinned this candidate, the resolver found nothing of
# its own, and the case study reported no holdout although a later candidate had
# one. `training_run_fitted_for_the_holdout` is the other half: a model fitted on
# the validation folds publishes over the holdout window, so split plus a non-null
# Sharpe does not make a row a holdout result, and it is not expressible in SQL.
carrier_row = db.execute(
"""
SELECT t.family, t.config_name, t.label,
p.checkpoint_value, p.checkpoint_kind
FROM prediction_sets p
JOIN training_runs t ON t.training_hash = p.training_hash
WHERE p.prediction_hash = ?
""",
(cand["prediction_hash"][i],),
).fetchone()
if carrier_row is None:
continue
probe_rows = db.execute(
f"""
SELECT t.spec_json FROM prediction_sets p
JOIN training_runs t ON p.training_hash = t.training_hash
JOIN backtest_runs b ON p.prediction_hash = b.prediction_hash
AND b.stage IN ('signal','allocation','risk_overlay','holdout')
JOIN backtest_metrics bm ON b.backtest_hash = bm.backtest_hash
WHERE {" AND ".join(ho_clauses)}
AND t.family = ?
AND t.config_name = ?
AND t.label = ?
AND p.checkpoint_value IS ?
AND p.checkpoint_kind IS ?
AND bm.sharpe IS NOT NULL
""",
ho_params + list(carrier_row),
).fetchall()
if any(training_run_fitted_for_the_holdout(probe[0]) for probe in probe_rows):
return {"spec": spec, "prediction_hash": cand["prediction_hash"][i]}
finally:
db.close()
return None
def _holdout_lineage_for(
cs: str,
leader_label: str,
strategy_spec: dict | None = None,
*,
label_restriction: frozenset[str] | None,
rung: dict | None,
prefer_prediction_hash: str | None = None,
retired_hashes: frozenset[str] | None = None,
) -> dict | None:
"""Return ``{backtest_hash, prediction_hash, family, config_name, label}``
for the highest-Sharpe holdout backtest registered in this case study,
honoring per-CS cluster restrictions but **not** the leader's label.
``leader_label`` is intentionally unused in the SQL — kept for call-site
symmetry with ``_val_backtest_for_lineage``. The holdout's *own* label is
returned so callers can pair it against matching benchmarks (the label may
differ from the validation rank-1 when ``generate_holdout``'s degeneracy
fallback accepts a candidate on a different label).
When ``strategy_spec`` is provided, the holdout pick is restricted to
backtests with the same full (signal, allocation, risk) tuple as val's
rank-1 carrier, so the val→holdout comparison stays apples-to-apples.
"""
case_dir = get_case_study_dir(cs)
db_path = case_dir / "run_log" / "registry.db"
if not db_path.exists():
return None
clauses = ["p.split = 'holdout'"]
params: list[object] = []
if label_restriction:
placeholders = ",".join("?" for _ in label_restriction)
clauses.append(f"t.label IN ({placeholders})")
params.extend(label_restriction)
if rung is not None:
clauses.append(
"COALESCE(json_extract(b.spec_json, '$.strategy.signal.universe_filter'), 'full') = ?"
)
params.append(rung["universe_filter"])
if rung["exit_at_max_days"] is None:
clauses.append(
"json_extract(b.spec_json, '$.strategy.signal.exit_at_max_days') IS NULL"
)
else:
clauses.append("json_extract(b.spec_json, '$.strategy.signal.exit_at_max_days') = ?")
params.append(rung["exit_at_max_days"])
if retired_hashes:
# ``_retired_prediction_hashes`` carries the recorded validation retirements across
# to the holdout rows that share a training run and checkpoint, so this filters on the
# identity the row actually has.
#
# It reaches a holdout scored from the same trained model. One produced by the
# canonical *retrain* registers its own training hash and shares no key with the
# validation identity that was retired - measured: no holdout training spec in any
# registry here references a validation identity, so there is nothing to join on.
#
# What prevents a retired carrier from reaching a holdout is upstream instead:
# `strategy_analysis.resolve_canonical_rank1_lineage` ranks over published members
# only, so a retrain created from here on descends from a live carrier by
# construction. Holdout rows written
# before that are stale artifacts of an earlier selection, and they are regenerated.
clauses.append("p.prediction_hash NOT IN (SELECT value FROM json_each(?))")
params.append(json.dumps(sorted(retired_hashes)))
spec_clauses, spec_params = _full_strategy_clauses(strategy_spec)
clauses.extend(spec_clauses)
params.extend(spec_params)
where_sql = " AND ".join(clauses)
db = sqlite3.connect(str(db_path))
db.row_factory = sqlite3.Row
try:
# Same-lineage preference: given the validation rank-1's own prediction
# set, prefer a holdout sharing its trained model AND its checkpoint.
# Both are read from that one row here rather than accepted as separate
# arguments, because a caller that passes the training hash and forgets
# the checkpoint reintroduces the defect while looking correct.
#
# The checkpoint has to be pinned because a trained model registers one
# prediction set per declared checkpoint and they share a strategy spec,
# so the training hash alone leaves one indistinguishable candidate per
# checkpoint. This branch writes the hash the ``val_rank1_self`` pair is
# stored under and ``select_holdout_self_backtest`` reads it back, so the
# two must agree. Ordering by ``backtest_hash`` rather than by Sharpe
# keeps the choice off holdout performance either way.
if prefer_prediction_hash is not None:
carrier = db.execute(
"""
SELECT t.family, t.config_name, t.label,
p.checkpoint_value, p.checkpoint_kind, t.spec_json
FROM prediction_sets p
JOIN training_runs t ON t.training_hash = p.training_hash
WHERE p.prediction_hash = ?
""",
(prefer_prediction_hash,),
).fetchone()
if carrier is not None:
carrier_spec_json = carrier["spec_json"]
carrier_key = list(carrier)[:5]
# Matched on the declared configuration, not on the carrier's training
# hash. A holdout prediction produced correctly carries a NEW training
# identity - the same configuration refitted on the holdout fold - so
# matching on the validation training hash finds only a holdout scored
# from the validation-fitted model. ``select_holdout_self_backtest``
# reads back the hash this branch writes the ``val_rank1_self`` pair
# under, so the two apply the same rule.
rows_ = db.execute(
f"""
SELECT t.family, t.config_name, t.label,
p.prediction_hash, b.backtest_hash, t.spec_json
FROM prediction_sets p
JOIN training_runs t ON p.training_hash = t.training_hash
JOIN backtest_runs b ON p.prediction_hash = b.prediction_hash
AND b.stage IN
('signal','allocation','risk_overlay','holdout')
JOIN backtest_metrics bm ON b.backtest_hash = bm.backtest_hash
WHERE {where_sql}
AND t.family = ?
AND t.config_name = ?
AND t.label = ?
AND p.checkpoint_value IS ?
AND p.checkpoint_kind IS ?
ORDER BY b.backtest_hash
""",
params + carrier_key,
).fetchall()
# Naming the carrier pins the configuration and the checkpoint, and that is
# normally one candidate. It is not guaranteed to be: one prediction set can
# carry several backtests - a replay under a different strategy spec, an
# experimental allocator sharing the holdout prediction - and they survive
# this filter together. Returning the first in `backtest_hash` order would
# decide on nothing the carrier determines, which is the same defect the
# unpinned branch below refuses, so it refuses here too rather than only
# where the caller happened not to pin.
# Fitted for the holdout AND a refit of this specification. The four
# columns the query filters on are a configuration's NAME, and a name is
# reused across generations - refit a study after its features change and the
# new runs carry the same family, config_name, label and checkpoint as the
# old ones. On the current registries fx_pairs has 144 configuration groups
# spanning more than one feature-artifact generation and etfs has 10, so
# without the second condition a holdout fitted on features the study no
# longer publishes can be the sole coarse match and get reported as the
# carrier's own holdout.
eligible = [
candidate
for candidate in rows_
if training_run_fitted_for_the_holdout(candidate["spec_json"])
and is_refit_of(candidate["spec_json"], carrier_spec_json)
]
if len({candidate["backtest_hash"] for candidate in eligible}) > 1:
raise ValueError(
f"{len(eligible)} holdout backtests match the pinned carrier "
f"{prefer_prediction_hash} for {cs} - same configuration, same "
"checkpoint, same strategy. Choosing between them would rank the "
"holdout on its own result. Retire the replays that are not this "
"study's holdout, so one candidate remains."
)
if eligible:
resolved = dict(eligible[0])
resolved.pop("spec_json")
return resolved
# A named carrier with no eligible holdout is an ANSWER, not a reason to look
# elsewhere. Falling through to the unpinned query below made the pin a mere
# preference: where the selected checkpoint had no holdout and a sibling
# checkpoint did, the fallback returned the sibling's, and the reader-facing
# table then reported a checkpoint that validation never selected. There is no
# weaker sense in which that is the carrier's holdout.
#
# `carrier is None` lands here too, and for the same reason: the caller named a
# prediction set the registry does not have, and "some other holdout" is not a
# better answer to that than none.
return None
# The caller named no carrier, so this is the unpinned fallback. Two filters, and
# neither is optional.
#
# Only runs actually refitted for the holdout are eligible: a model fitted on the
# validation folds can publish predictions over the holdout window, and it is not a
# holdout result whatever its Sharpe.
#
# And what survives has to be ONE candidate. Anything that reaches here cannot be
# separated on what the registry records, so picking among them means picking by
# holdout Sharpe - choosing the evaluation by its own result, which is the one thing
# this module must never do. The refusal is the answer; the caller resolves it by
# naming the validation carrier, not by this function guessing.
rows = db.execute(
f"""
SELECT DISTINCT t.family, t.config_name, t.label,
p.prediction_hash, b.backtest_hash, p.training_hash, t.spec_json
FROM prediction_sets p
JOIN training_runs t ON p.training_hash = t.training_hash
JOIN backtest_runs b ON p.prediction_hash = b.prediction_hash
AND b.stage IN ('signal','allocation','risk_overlay','holdout')
JOIN backtest_metrics bm ON b.backtest_hash = bm.backtest_hash
WHERE {where_sql}
ORDER BY b.backtest_hash
""",
params,
).fetchall()
finally:
db.close()
rows = [row for row in rows if training_run_fitted_for_the_holdout(row["spec_json"])]
if not rows:
return None
# Ambiguity is ambiguity however it arises, and it arises three ways: several trained
# models, several checkpoints of one trained model (one prediction set each, sharing a
# strategy spec), and several backtests hanging off one prediction set. Only the first
# used to refuse; the other two fell through to `rows[0]` under `ORDER BY b.backtest_hash`,
# which is arbitrary with respect to the configuration and therefore picks on the holdout's
# own result as surely as ordering by Sharpe would.
if len({row["backtest_hash"] for row in rows}) > 1:
lineages = {row["training_hash"] for row in rows}
raise ValueError(
f"{len(rows)} holdout backtests across {len(lineages)} trained model(s) were "
f"refitted for the holdout and match the carrier spec for {cs}. Choosing between "
"them would rank the holdout on its own result. Pass the validation rank-1's "
"prediction hash as `prefer_prediction_hash`, which pins the configuration and the "
"checkpoint, so the holdout is resolved from the carrier that was selected rather "
"than from the holdout scores."
)
row = rows[0]
return {
k: row[k] for k in ("family", "config_name", "label", "prediction_hash", "backtest_hash")
}
def _val_backtest_for_lineage(
cs: str,
family: str,
config_name: str,
label: str,
*,
prediction_hashes: list[str] | None = None,
) -> str | None:
"""Return the highest-Sharpe validation signal-stage backtest_hash for the
given (family, config_name, label) lineage, or None if absent.
Used by ``val_rank1_self`` pair construction so the comparison stays
*within* a lineage when the holdout retrain came from a fallback candidate.
"""
case_dir = get_case_study_dir(cs)
db_path = case_dir / "run_log" / "registry.db"
if not db_path.exists():
return None
db = sqlite3.connect(str(db_path))
try:
row = db.execute(
"""
SELECT b.backtest_hash
FROM prediction_sets p
JOIN training_runs t ON p.training_hash = t.training_hash
JOIN backtest_runs b ON p.prediction_hash = b.prediction_hash
AND b.stage = 'signal'
JOIN backtest_metrics bm ON b.backtest_hash = bm.backtest_hash
WHERE p.split = 'validation'
AND t.family = ?
AND t.config_name = ?
AND t.label = ?
{scope}
ORDER BY bm.sharpe DESC NULLS LAST
LIMIT 1
""".format(
scope=(
" AND p.prediction_hash IN (SELECT value FROM json_each(?))"
if prediction_hashes is not None
else ""
)
),
(family, config_name, label)
+ ((json.dumps(list(prediction_hashes)),) if prediction_hashes is not None else ()),
).fetchone()
finally:
db.close()
return row[0] if row else None
def _populate_pair(
cs,
challenger_hash,
benchmark_hash,
benchmark_kind,
challenger_returns,
benchmark_returns,
ppy,
label,
*,
disjoint_windows: bool = False,
challenger_overlays_baseline: bool = False,
benchmark_label: str | None = None,
write_case_dir: Path | None = None,
):
"""Compute and register one paired-metric row. Idempotent UPSERT.
With ``disjoint_windows=True`` (val→holdout decay), each side is bootstrapped
over its full window and the difference distribution is built from those draws,
because two windows that share no timestamps leave no difference series to
resample. That is the absence of a pairing, not independence: see
:func:`case_studies.utils.uncertainty.compute_independent_diff_uncertainty` for
what the resulting interval does and does not cover. Otherwise, the streams are
inner-joined on timestamp and a paired stationary bootstrap runs on the aligned
diff series.
``challenger_overlays_baseline`` says what a leading flat run on the challenger
means, and the two pair shapes here answer differently. Against the equal-weight
benchmark the challenger is an independent strategy whose returns begin at the
first bar rather than at its first signal, so those rows are warmup and the
default drops them. Across a stage transition the challenger is built on top of
the baseline and both are live from the same session, so a flat challenger there
is a position it chose to hold and the comparison keeps it. See
:func:`case_studies.utils.uncertainty.joint_returns`.
"""
min_n = _min_paired_n(ppy)
if disjoint_windows:
c_arr = challenger_returns.sort("timestamp")["ret"].to_numpy()
b_arr = benchmark_returns.sort("timestamp")["ret"].to_numpy()
finite_c = np.isfinite(c_arr)
finite_b = np.isfinite(b_arr)
c_arr, b_arr = c_arr[finite_c], b_arr[finite_b]
if c_arr.size < min_n or b_arr.size < min_n:
return {
"cs": cs,
"kind": benchmark_kind,
"label": label,
"benchmark_label": benchmark_label if benchmark_label is not None else label,
"skip": f"insufficient_disjoint:n_c={c_arr.size},n_b={b_arr.size}",
}
n_overlap = min(c_arr.size, b_arr.size)
paired = compute_independent_diff_uncertainty(
c_arr,
b_arr,
periods_per_year=ppy,
case_study=cs,
label=label,
n_boot=2000,
seed=42,
)
else:
aligned = challenger_returns.join(
benchmark_returns, on="timestamp", how="inner", suffix="_b"
)
if aligned.height < min_n:
return {
"cs": cs,
"kind": benchmark_kind,
"label": label,
"benchmark_label": benchmark_label if benchmark_label is not None else label,
"skip": f"insufficient_overlap:n={aligned.height}",
}
c_arr = aligned["ret"].to_numpy()
b_arr = aligned["ret_b"].to_numpy()
# Measured here, applied once inside `compute_paired_uncertainty`, which trims
# whatever it is handed. Both rules happen to survive being applied twice, so passing
# the trimmed pair on would work today; it would stop working silently the moment a
# rule stops being idempotent, and the figure registered as `n_overlap` would then
# name a sample the bootstrap never ran on.
n_overlap = joint_returns(
c_arr, b_arr, challenger_overlays_baseline=challenger_overlays_baseline
)[0].size
if n_overlap < min_n:
return {
"cs": cs,
"kind": benchmark_kind,
"label": label,
"benchmark_label": benchmark_label if benchmark_label is not None else label,
"skip": f"insufficient_after_coerce:n={n_overlap}",
}
paired = compute_paired_uncertainty(
c_arr,
b_arr,
periods_per_year=ppy,
case_study=cs,
label=label,
n_boot=2000,
seed=42,
challenger_overlays_baseline=challenger_overlays_baseline,
)
if not paired:
return {
"cs": cs,
"kind": benchmark_kind,
"label": label,
"benchmark_label": benchmark_label if benchmark_label is not None else label,
"skip": "uncertainty_empty",
}
register_paired_metrics(
cs,
challenger_hash,
benchmark_hash,
paired,
benchmark_kind=benchmark_kind,
periods_per_year=ppy,
case_dir=write_case_dir,
)
# disjoint path: paired carries n_c/n_b (post-coerce per-side sizes); use
# min so n_overlap reflects what the bootstrap actually used. paired path:
# no n_c/n_b, n_overlap already the post-`joint_returns` length.
n_actual = n_overlap
n_c = paired.get("n_c")
n_b = paired.get("n_b")
if n_c is not None and n_b is not None:
n_actual = int(min(float(n_c), float(n_b)))
return {
"cs": cs,
"kind": benchmark_kind,
"label": label,
"benchmark_label": benchmark_label if benchmark_label is not None else label,
"n_overlap": n_actual,
"sharpe_diff": paired.get("sharpe_diff"),
"sharpe_diff_ci_lo": paired.get("sharpe_diff_ci95_lo"),
"sharpe_diff_ci_hi": paired.get("sharpe_diff_ci95_hi"),
"info_ratio": paired.get("info_ratio"),
"p_value": paired.get("p_value"),
}
def _drop_retired_generations(cs: str, cand):
"""Candidates whose own publisher still publishes them.
Every ranking in this module sorts on `sharpe` over whatever the registry holds, and a
superseded generation is still complete, still `current` under its schema version, and
still ranks. Passing a resolved `carrier` fixes only the pairs that consult it; Pair #1
ranks for itself, so the retired row won there and the validation-side pair was written
against a backtest the case study no longer publishes - measured on fx_pairs, where the
live carrier then had no challenger row at all and the strategy-analysis notebook refused
for want of evidence that had been written under the retired hash.
Both sides are filtered, because a retired generation reaches a ranking through either.
The prediction side is the one that hides: a refit that changes no numbers publishes
identical predictions under a new identity, so old and new carry the same Sharpe to the
last digit and the sort returns whichever it likes. It goes through
``_retired_prediction_hashes`` rather than the recorded set, so a holdout prediction is
dropped along with the validation generation it was retrained from.
"""
from case_studies.research.population import superseded_members_at
if cand is None or cand.is_empty():
return cand
if "backtest_hash" in cand.columns:
retired = superseded_members_at(get_case_study_dir(cs), member_kind="backtest")
if retired:
cand = cand.filter(~pl.col("backtest_hash").is_in(list(retired)))
if "prediction_hash" in cand.columns:
retired_predictions = _retired_prediction_hashes(cs)
if retired_predictions:
cand = cand.filter(~pl.col("prediction_hash").is_in(list(retired_predictions)))
return cand
def _drop_degenerate_predictions(cs: str, cand):
"""Candidates whose prediction set selection refuses to consider.
A LASSO or ElasticNet fit that shrinks every coefficient to zero on a fold predicts a
constant there, so that fold ranks nothing and the pooled IC computed over it is biased
rather than a model result. ``degenerate_prediction_sql`` states the rule - both limbs of
it, the NULL IC of an all-tied fold and the denormal one of a fold constant only to display
precision - and
``selectable_validation_candidates`` applies it, which is why the published carrier cannot
be one of these.
A pair is the other publication path and had no such filter. The sweep backtests the whole
declared population rather than a shortlist, so the registry does hold backtests on
degenerate sets: measured on us_equities_panel 2026-09-14, 15 of its prediction sets are
degenerate, four signal-stage backtests stand on two of them, and both sat at Sharpe 0.6062
against a 0.8977 leader - third and fourth in the stage, so ranking alone did not catch it
and gives no reason to expect it to as the sweep continues.
Applied wherever ``_drop_retired_generations`` is, for the same reason: every ranking in
this module sorts on ``sharpe`` over whatever the registry holds.
"""
if cand is None or cand.is_empty() or "prediction_hash" not in cand.columns:
return cand
degenerate = degenerate_prediction_hashes(get_case_study_dir(cs))
if not degenerate:
return cand
return cand.filter(~pl.col("prediction_hash").is_in(list(degenerate)))
def _eligible_candidates(
cs: str, cand, *, label_restriction: frozenset[str] | None, rung: dict | None
):
"""Every filter a ranking in this module owes its candidate pool, in one place.
The three rankings here - the carrier walk, pair #1, and the no-carrier leader - had this
chain written out three times, which is how the degeneracy filter came to be missing from
all of them while selection had it: a filter added to one copy is not added to the others,
and nothing reads as wrong at any single site.
Retirement and degeneracy are both facts about whether the row may be published at all.
The benchmark exclusion is about what a challenger is. The label and rung restrictions are
the case study's own scope. What is NOT here is the carrier pin, which pair #1 deliberately
does not apply.
"""
cand = _drop_retired_generations(cs, cand)
cand = _drop_degenerate_predictions(cs, cand)
if cand is None or cand.is_empty():
return cand
if "family" in cand.columns:
cand = cand.filter(pl.col("family") != "benchmark")
if label_restriction and "label" in cand.columns:
cand = cand.filter(pl.col("label").is_in(list(label_restriction)))
return _apply_rung_restriction(cand, rung)
def populate_paired_metrics(
cs: str,
explorer: BacktestExplorer | None = None,
*,
label_restriction: frozenset[str] | None = None,
rung: dict | None = None,
carrier: CarrierScope,
periods_per_year: int | None = None,
verbose: bool = True,
replace_all: bool,
write_case_dir: Path | None = None,
prediction_hashes: PredictionScope,
) -> list[dict]:
"""Compute all six paired-bootstrap pair types for ``cs`` and register them.
Extracted from the two producer loops in
``20_strategy_synthesis/01_aggregate_synthesis.py``, specialized to a single
case study. The per-CS selection config that Ch20 reads from module globals
is passed in:
* ``label_restriction`` - ``strategy_analysis.LABEL_RESTRICTIONS.get(cs)`` (e.g.
sp500_options → ``frozenset({'ret_to_expiry'})``); None for most CSs.
* ``rung`` - ``{"predicate", "universe_filter", "exit_at_max_days"}``, plus
``label`` where the pin names one, for the rung-pinned CSs (sp500_options,
nasdaq100_microstructure); None else. A line naming
``us_firm_characteristics -> config_name == 'default_huber'`` stood here
until 2026-09-14, dangling under this bullet and describing a pin deleted on
2026-08-25 for selecting that case study's weakest advanced configuration
(`20_strategy_synthesis/01_aggregate_synthesis.py:365`).
* ``periods_per_year`` — the annualization factor. Defaults to the case
study's own ``evaluation.periods_per_year`` declaration rather than to a
cadence, so a caller that omits it gets its own scale instead of someone
else's.
* ``carrier`` — a ``resolve_canonical_rank1_lineage`` result, or ``NO_CARRIER``.
With a lineage, pairs #2-6 use its validation and holdout backtests instead of
re-ranking the registry here, and pair #1 is pinned to its validation backtest
too - the code below refuses rather than ranking when a carrier is supplied,
because a pair #1 registered under a backtest the case study does not report
leaves its carrier with no validation-to-benchmark evidence. This bullet said
pair #1 was unaffected until 2026-09-14, which had not been true since that
refusal landed. ``NO_CARRIER`` keeps the legacy ranking, which is not
the canonical selection - it orders on raw Sharpe and applies neither the
common-support re-ranking nor the restrictions the resolver holds - so a caller
that can resolve the lineage should pass it. The rung-pinned case studies
(sp500_options, nasdaq100_microstructure) restrict on a dimension the resolver
does not know, which is why the legacy ranking still exists at all.
``replace_all`` makes the call a complete snapshot: pairs it did not write are
deleted, so a rebuild under a different selection does not leave the previous
selection's rows behind. Registration alone is an UPSERT keyed on
``(challenger_hash, benchmark_hash)``, which cannot remove a row it no longer
produces, so False is additive: the previous selection's rows survive alongside
the corrected ones. A call that writes nothing prunes nothing - that is a failed
rebuild, not an empty snapshot.
``periods_per_year`` used to be ``freq: str = "daily"``, resolved through a
name-to-count map. That default is silently right for the six case studies that
annualize at 252 and silently wrong for the rest: ``us_firm_characteristics`` is
monthly, so every Sharpe difference and interval it wrote was scaled by sqrt(252)
rather than sqrt(12), a factor of 4.58 on numbers a notebook prints as its holdout
closure. Worse, ``_min_paired_n(252)`` returns 21, so the twelve observation
holdout pairs were skipped and the table was missing rows with nothing recording
the omission.
An integer rather than a cadence name because the name was only ever converted
back to a number, and the caller holds the number. Reading the declaration by
default is what stops the next non-252 case study inheriting the wrong scale by
saying nothing.
Returns the list of per-pair summary dicts (mirrors the ``paired_rows`` +
``extra_paired_rows`` the Ch20 producer builds); each pair is also written to
``backtest_paired_metrics`` via ``register_paired_metrics``.
``prediction_hashes`` restricts every candidate read to that population, or
``ENTIRE_REGISTRY`` for the whole-registry read. A pair is a comparison between two
strategies the caller reports; selecting either side from the whole registry lets a
retired generation be the challenger or the benchmark, and the difference is then
measured against a strategy its own publisher replaced. Detecting those rows and
rebuilding without this would write them back unchanged.
``carrier``, ``replace_all`` and ``prediction_hashes`` carry no default. Each decides
what the numbers this writes are computed over, and each used to default to the widest
reading, so omitting one type-checked, ran, and produced plausible rows that were wrong
exactly when the registry held something the caller does not report - invisible in
review, in CI and in the output. Requiring them costs one line per call site and makes
the wide readers greppable. See ``uncertainty.ENTIRE_REGISTRY``.
``write_case_dir`` redirects the registry *write* to an alternate case dir
(reads still come from the live tree) — used by the verification harness to
write into a temp registry copy non-destructively.
"""
if explorer is None:
explorer = BacktestExplorer(cs)
# Both scopes are stated by the caller and carry no default; the sentinels are
# normalized here so the rest of the body reads the same as it did when they were
# `None`. `ENTIRE_REGISTRY` is the whole-registry read, `NO_CARRIER` the raw-Sharpe
# re-rank. See `uncertainty.ENTIRE_REGISTRY` for why neither is a default.
live = None if isinstance(prediction_hashes, EntireRegistry) else list(prediction_hashes)
lineage = None if isinstance(carrier, NoCarrier) else carrier
if periods_per_year is None:
from case_studies.utils.uncertainty import periods_per_year_from_setup
periods_per_year = int(periods_per_year_from_setup(cs))
ppy = int(periods_per_year)
rows: list[dict] = []
written_keys: set[tuple[str, str]] = set()
def _pair(cs_, challenger_hash, benchmark_hash, *args, **kwargs):
result = _populate_pair(cs_, challenger_hash, benchmark_hash, *args, **kwargs)
if "skip" not in result:
written_keys.add((challenger_hash, benchmark_hash))
return result
# -- Pair #1: signal rank-1 (overall) ↔ equal-weight (overall) -----------
cand = pl.concat(
[_best_for_rung(explorer, s, rung, prediction_hashes=live) for s in _PAIRED_STAGES],
how="diagonal_relaxed",
)
skip_pair1 = False
if cand.is_empty() or "backtest_hash" not in cand.columns:
skip_pair1 = True
if not skip_pair1:
# NB: pair #1 (Ch20 Loop A) applies ONLY the rung restriction — no
# carrier pin — unlike pairs #2-6 (Loop B), which apply both. Preserve
# that asymmetry so carrier-pinned CSs (us_firm_characteristics) match.
cand = _eligible_candidates(cs, cand, label_restriction=label_restriction, rung=rung)
if cand.is_empty():
skip_pair1 = True
if not skip_pair1:
# This ranking is on raw Sharpe, and the canonical resolver is not: when a conformal
# candidate is in the field it re-ranks everything on exact common timestamp support,
# so it can return a lower raw-Sharpe row. It also happens that the two tie exactly -
# a risk overlay that never binds produces the same returns as the allocation stage
# under it, to the last digit - and the dedupe then keeps whichever the sort emitted.
#
# Either way the pair ends up registered under a backtest the case study does not
# report, and a notebook asking for its carrier's validation-to-benchmark evidence
# finds none. So a supplied carrier is used rather than ranked against: the caller
# resolved it through the canonical selection, which is the answer this ranking is a
# cheaper approximation of. With no carrier the sort stands, with `backtest_hash` as
# a final key so the choice is at least deterministic.
carrier_backtest = str(lineage["val_backtest_hash"]) if lineage else None
cand1 = cand.sort(["sharpe", "backtest_hash"], descending=[True, False]).unique(
subset=["prediction_hash"], keep="first", maintain_order=True
)
if carrier_backtest is not None:
pinned = cand.filter(pl.col("backtest_hash") == carrier_backtest)
if pinned.is_empty():
raise RuntimeError(
f"the carrier {carrier_backtest} passed for {cs} is not among the "
f"{cand.height} candidates this ranking sees. Pair #1 would be registered "
"under a different backtest than the one the case study reports."
)
cand1 = pinned
leader_hash = cand1["backtest_hash"][0]
leader_label = cand1["label"][0] if "label" in cand1.columns else None
if leader_label:
bench_resolution = _benchmark_returns_from_artifact(cs, leader_label)
chal = _aligned_returns(cs, leader_hash)
if bench_resolution and chal is not None:
benchmark_hash, base, resolved_bench_label = bench_resolution
min_n = _min_paired_n(ppy)
aligned = chal.join(base, on="timestamp", how="inner", suffix="_b")
if aligned.height >= min_n:
c_arr = aligned["ret"].to_numpy()
b_arr = aligned["ret_b"].to_numpy()
# Sized here, trimmed once inside `compute_paired_uncertainty`; see
# `_populate_pair` for why the pair is not trimmed on the way in.
if joint_returns(c_arr, b_arr)[0].size >= min_n:
paired = compute_paired_uncertainty(
c_arr,
b_arr,
periods_per_year=ppy,
case_study=cs,
label=leader_label,
n_boot=2000,
seed=42,
)
if paired:
benchmark_kind = (
f"{SIGNAL_BASELINE_BY_CASE_STUDY.get(cs, 'equal_weight')}"
"_side_artifact"
)
register_paired_metrics(
cs,
leader_hash,
benchmark_hash,
paired,
benchmark_kind=benchmark_kind,
periods_per_year=ppy,
case_dir=write_case_dir,
)
written_keys.add((leader_hash, benchmark_hash))
rows.append(
{
"case_study": DISPLAY_NAMES.get(cs, cs),
"kind": benchmark_kind,
"label": leader_label,
"benchmark_label": resolved_bench_label,
"sharpe_diff": paired.get("sharpe_diff"),
"sharpe_diff_ci_lo": paired.get("sharpe_diff_ci95_lo"),
"sharpماخذ کا حوالہ دیتے ہوئے مکمل متن دکھایا گیا ہے، ماخذ کے لائسنس کے تحت۔ لائسنس: MIT
یہ خلاصہ اصل ماخذ سے Stratmill کے تحقیقی ایجنٹ نے لکھا ہے؛ یہ ماخذ کی نقل نہیں۔