Оценка бэктестов по характеристикам компаний US с парной оценкой неопределённости и стресс-тестами издержек
Сводка
В аналитической записной книжке объясняется, как оценивать стратегию на акциях US, построенную на характеристиках компаний, используя результаты реестра, а не новые бэктесты. Рассматривается статистика доходности с учётом неопределённости, конфигурации ранжируются на общей выборке, а парные сравнения используются для оценки выбранной стратегии относительно бенчмарка с равными весами. Отложенная выборка определяется по полной спецификации выбранного запуска на валидации, чтобы не выбирать стратегию по результатам отложенного периода.
Также изучаются сложности реализации с помощью расширенной сетки издержек, рассчитанной на низкую ликвидность выбранных акций, и приводится скорректированная на отбор статистика, например варианты дефлированного коэффициента Шарпа для семейства и когорты, а также вероятность переобучения на бэктесте. Описаны пороги по коэффициенту Шарпа на валидации и сравнению с отложенной выборкой; подчёркивается, что единственного годового отложенного периода недостаточно для точных выводов о снижении эффективности. Выводы относятся только к основной месячной метке доходности; заявленные варианты не тестировались на бэктесте. Короткие позиции предполагают доступность займа по ставкам, заданным в исследовании, а издержки могут существенно уменьшить разницу между длинной и короткой сторонами.
Ключевые идеи
- Используйте доверительные интервалы и парные показатели для оценки бэктестов, а не полагайтесь только на точечные оценки.
- Определяйте оценку на отложенной выборке по спецификации стратегии, выбранной на валидации, чтобы избежать отбора по отложенным данным.
- Проверяйте устойчивость к издержкам по сетке, учитывающей низкую ликвидность выбранных акций микрокапитализации.
- Считайте интервал на отложенной выборке за один год, охватывающий ноль, ограниченным свидетельством из-за короткого периода наблюдений.
- Интерпретируйте результаты с учётом основной метки и заданных предположений о ставках займа для коротких позиций.
Теги
Полный текст
# US Firm Characteristics - Strategy Analysis
# US Firm Characteristics - Strategy Analysis
This notebook converts the backtest registry for the US
firm-characteristics case study into a per-case-study strategy
assessment. Metrics are reported with a block-bootstrap
confidence interval, every comparison goes through
`backtest_paired_metrics`, and the holdout closure uses paired rather
than point-difference reasoning. Cross-case-study comparison is reserved
for Chapter 20.
**Learning objectives**
- Read uncertainty-aware backtest metrics (Sharpe with CI, PSR, DSR) from
the registry rather than transcribing point estimates.
- Resolve the single holdout evaluation from the selected validation run's
complete strategy specification rather than selecting on holdout performance.
- Quantify implementation friction with a micro-cap-realistic extended
cost grid driven by the illiquidity profile of selected stocks.
- Distinguish a decay whose CI includes zero because the effect is absent
from one whose CI includes zero because a one-year holdout is short.
**Book reference**: Chapter 20, §20.1 (the §9 handoff feeds Ch20's
cross-case-study aggregation).
**Prerequisites**: case-study pipeline through
[`16_holdout_backtest`](16_holdout_backtest.ipynb); the registry
(`case_studies/us_firm_characteristics/run_log/registry.db`).
**Scope**: no training or re-backtesting. The case-study pipeline registers
the derived paired-metrics rows before this notebook runs, after which the
analysis is registry-read only.
```python
"""US Firm Characteristics - Strategy Analysis."""
import datetime as _dt
import json
import sqlite3
import matplotlib.dates as mdates
import matplotlib.pyplot as plt
import numpy as np
import pandas as pd
import polars as pl
# ml4t.diagnostic loads cudart; torch must import first, so its bundled runtime wins symbol
# resolution. Imported for that side effect alone, which ruff cannot see - without the noqa
# a dead-import sweep deletes it and the notebook fails on the cudart load.
import torch # noqa: F401
import yaml
from matplotlib.colors import LinearSegmentedColormap
```
```python
from ml4t.diagnostic.evaluation import PortfolioAnalysis
from ml4t.diagnostic.integration import (
BacktestReportMetadata,
generate_tearsheet_from_run_artifacts,
)
from ml4t.diagnostic.visualization.backtest.tail_risk import plot_tail_risk_analysis
from ml4t.diagnostic.visualization.portfolio.drawdown_plots import (
plot_drawdown_underwater,
)
from ml4t.diagnostic.visualization.portfolio.returns_plots import (
plot_annual_returns_bar,
plot_cumulative_returns,
plot_monthly_returns_heatmap,
plot_returns_distribution,
)
from ml4t.diagnostic.visualization.portfolio.risk_plots import plot_rolling_sharpe
```
```python
from case_studies.research.population import split_unpublished_members
from case_studies.research.workspace import open_study
from case_studies.utils.backtest_explorer import BacktestExplorer
from case_studies.utils.benchmark import load_benchmark_metrics, load_benchmark_returns
from case_studies.utils.conformal import CALIBRATION_VERSION
from case_studies.utils.external_benchmarks import (
align_to_strategy_monthly as align_external_benchmark_monthly,
)
from case_studies.utils.external_benchmarks import (
compute_benchmark_diagnostics,
compute_subperiod_diagnostics,
load_ff_market_returns,
)
```
```python
from case_studies.utils.factor_attribution import (
compute_bootstrap_ci,
compute_rolling_exposures,
load_factor_data,
plot_attribution_waterfall,
plot_rolling_exposures,
run_factor_regression,
run_placebo_benchmark,
)
from case_studies.utils.notebook_contracts import degenerate_prediction_sql
from case_studies.utils.registry import (
load_backtest_fold_metrics,
load_backtest_metrics,
load_paired_metrics,
load_prediction_index,
)
from case_studies.utils.strategy_analysis import (
ci_status,
fmt_gate,
gate1_validation_sharpe_geq_zero,
gate2_holdout_diff_not_excludes_zero_negatively,
rank_backtests_on_common_support,
resolve_solvent_carrier,
select_holdout_self_backtest,
)
from utils.paths import display_path, get_case_study_dir, get_output_dir
from utils.reproducibility import set_global_seeds
from utils.style import COLORS, add_message_title, ml4t_diverging, show_with_alt
```
Papermill overrides names in the cell below and nowhere else. Without it the
notebook cannot be run at reduced scale through the same runner production uses,
which is what the pre-production check requires.
```python
CASE_STUDY = "us_firm_characteristics"
SEED = 42
# Both names stay bound here although nothing below reads them: that is what makes the harness
# force preview and supply a workspace - `_declares_tier_and_workspace` in `tests/pm_helpers.py`
# looks for exactly this pair. Without them the canonical branch regenerates in place, which
# needs symlinks a CI checkout does not have.
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""
```
The study is opened before any path or registry read. Under the preview tier, opening it
activates a workspace and rewrites `ML4T_OUTPUT_DIR` process-wide; a `CASE_DIR` or a
`BacktestExplorer` built first would address the released registry while everything after it
reads the preview one. It is opened once and reused - a second `open_study` further down
re-activates, and where the two calls disagree about the tier the notebook silently changes
registry mid-page.
```python
study = open_study(CASE_STUDY, execution_tier=EXECUTION_TIER, workspace=WORKSPACE or None)
```
```python
PRIMARY_LABEL = "fwd_ret_1m" # setup.yaml primary; benchmark series keyed here
PERIODS_PER_YEAR = 12 # monthly cadence
CASE_DIR = get_case_study_dir(CASE_STUDY)
OUTPUT_DIR = get_output_dir(20, CASE_STUDY)
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
set_global_seeds(SEED)
with open(CASE_DIR / "config" / "setup.yaml") as f:
setup = yaml.safe_load(f)
explorer = BacktestExplorer(CASE_STUDY)
print(explorer)
```
Compact formatters keep unavailable registry values explicit in the
reader-facing tables.
```python
def _fmt_ci(point: float | None, lo: float | None, hi: float | None, fmt: str = ".3f") -> str:
"""Compact `point [lo, hi]` formatter with NULL-safety."""
if point is None:
return "-"
p = format(point, fmt)
if lo is None or hi is None:
return f"{p} [-, -]"
return f"{p} [{format(lo, fmt)}, {format(hi, fmt)}]"
```
```python
def _fmt(val: float | None, fmt: str = ".4f") -> str:
return "-" if val is None else format(val, fmt)
```
### The two derived tables this notebook reads
`populate_paired_metrics` is passed this notebook's own `PERIODS_PER_YEAR`. Its default
reads the case study's declaration and resolves to the same 12, so the argument states at
the call site what the function would otherwise resolve on its own. A monthly case study
annualizing at the square root of 252 is the defect this call site carried, and the
explicit value is what makes the cadence visible where the call is read.
`cohort_metrics` holds the selection-bias statistics §3 deflates by, and
`backtest_paired_metrics` holds the bootstrapped comparisons §6 closes the
holdout with. Both are derived from rows the sweep already registered rather
than from any new backtest, and both are computed here when they are absent so
that a registry regenerated by a reader carries them.
The guard matters more than the population. Both computations are idempotent and
reproduce the stored values exactly, but running them still rewrites the registry
file, and the registry is what every published number in this case study is read
from. So they run only when something this notebook needs is absent.
What "absent" means is the closure §6 actually reads: a `val_rank1_self` pair whose
challenger is the holdout backtest this notebook resolves. Neither weaker predicate
works, and both were tried here. A row count is non-empty forever once the validation
pairs are written, so the holdout pairs are never produced. The set of benchmark kinds
present is non-empty forever once ANY holdout has been paired - including a superseded
one - so re-running the holdout on a corrected configuration leaves the table carrying
the previous generation's pairs and §6 raises `Incomplete holdout closure pairs` while
the guard declines to repopulate. That is exactly what happened when this case study's
holdout was regenerated.
```python
from case_studies.utils.cohort_metrics import compute_and_register
from case_studies.utils.paired_metrics import populate_paired_metrics
from case_studies.utils.uncertainty import ENTIRE_REGISTRY
_db = CASE_DIR / "run_log" / "registry.db"
# The prediction sets their publishers still stand behind. `compute_and_register` scopes cohorts
# by prediction, and unscoped it computes them over the whole registry: 22 configurations here
# carry more than one training generation, and the pool grew by roughly 2,300 backtests when
# fwd_ret_1m_win and fwd_class_1m were backtested. K is what the deflated Sharpe divides by, so
# an unscoped call reports a correction computed over a population this page does not describe.
# Every declared label stays in - they compete at the baseline, and the selection ranges over
# all of them - so what is excluded is superseded generations, not variant labels.
_live_index = split_unpublished_members(
study,
load_prediction_index(CASE_STUDY, split="validation"),
)
LIVE_PREDICTIONS = _live_index.live["prediction_hash"].to_list()
if not LIVE_PREDICTIONS:
raise RuntimeError(f"no live validation prediction sets for {CASE_STUDY}")
print(f"Live prediction sets: {len(LIVE_PREDICTIONS):,}")
# Resolved before the guard because the guard asks about this holdout, not about holdouts
# in general. §6 re-resolves it and checks the two agree.
# `resolve_solvent_carrier` rather than the bare lineage resolver: same selection, and it
# additionally refuses a selected configuration whose equity reached zero.
_lineage = resolve_solvent_carrier(CASE_STUDY)
_expected_holdout = _lineage["holdout_backtest_hash"]
with sqlite3.connect(str(_db)) as _con:
_tables = {r[0] for r in _con.execute("SELECT name FROM sqlite_master WHERE type='table'")}
_n_cohorts = (
_con.execute("SELECT count(*) FROM cohort_metrics").fetchone()[0]
if "cohort_metrics" in _tables
else 0
)
_n_pairs = (
_con.execute("SELECT count(*) FROM backtest_paired_metrics").fetchone()[0]
if "backtest_paired_metrics" in _tables
else 0
)
_kinds_present = (
{
row[0]
for row in _con.execute(
"SELECT DISTINCT benchmark_kind FROM backtest_paired_metrics "
"WHERE challenger_hash IS ?",
(_expected_holdout,),
)
}
if "backtest_paired_metrics" in _tables
else set()
)
# Presence says the tables were built; it does not say they were built over the population
# this notebook reports. Scoping the cohorts to the live predictions changes K, and the
# previous unscoped rows survive under the same enum values - so a presence-only guard
# accepts them, prints "already present, nothing written", and the correction stays the one
# computed over a population the page does not describe. A cohort is stale when its
# membership is unrecorded, when its leader is a retired prediction, or when its stage still
# holds one; a pair is stale when either side is. Same test as `_stale_derived_rows` in
# `case_studies/etfs/20_strategy_analysis.py`, which reached it first.
_live_payload = json.dumps(LIVE_PREDICTIONS)
_stale_cohorts = (
_con.execute(
"""
SELECT COUNT(*) FROM cohort_metrics cm
WHERE cm.member_digest IS NULL
OR cm.leader_hash IN (
SELECT backtest_hash FROM backtest_runs
WHERE prediction_hash NOT IN (SELECT value FROM json_each(?1))
)
OR EXISTS (
SELECT 1 FROM backtest_runs r
WHERE r.stage IS cm.stage
AND r.prediction_hash NOT IN (SELECT value FROM json_each(?1))
)
""",
(_live_payload,),
).fetchone()[0]
if "cohort_metrics" in _tables
else 0
)
_stale_pairs = (
_con.execute(
"""
SELECT COUNT(*) FROM backtest_paired_metrics pm
WHERE pm.challenger_hash IN (
SELECT backtest_hash FROM backtest_runs
WHERE prediction_hash NOT IN (SELECT value FROM json_each(?1))
)
OR pm.benchmark_hash IN (
SELECT backtest_hash FROM backtest_runs
WHERE prediction_hash NOT IN (SELECT value FROM json_each(?1))
)
""",
(_live_payload,),
).fetchone()[0]
if "backtest_paired_metrics" in _tables
else 0
)
REQUIRED_PAIR_KINDS = {"val_rank1_self", "equal_weight_holdout_side_artifact"}
if (
_n_cohorts == 0
or not REQUIRED_PAIR_KINDS.issubset(_kinds_present)
or _stale_cohorts
or _stale_pairs
):
_cohort_counts = compute_and_register(CASE_STUDY, prediction_hashes=LIVE_PREDICTIONS)
# The lineage is passed rather than re-derived inside. Left to itself the populator
# ranks the registry on raw Sharpe, which is a fourth selector beside the resolver,
# this notebook and the costs sweep - and here it picked the retired conformal
# generation, so the pairs described a selected configuration the case study does not report.
# The cohort call above is scoped to `LIVE_PREDICTIONS` and this one is not: the
# pairs are selected from every registered prediction set. Stated rather than
# defaulted; narrowing it would change published numbers, so it is a separate
# decision. `replace_all=False` keeps the write additive, which
# is what this notebook has always done.
_paired_rows = populate_paired_metrics(
CASE_STUDY,
periods_per_year=PERIODS_PER_YEAR,
carrier=_lineage,
prediction_hashes=ENTIRE_REGISTRY,
replace_all=False,
)
_n_cohorts = sum(_cohort_counts[k] for k in ("family", "stagelabel", "label"))
_n_pairs = sum(1 for row in _paired_rows if "skip" not in row)
print(f"cohort_metrics: {_n_cohorts} rows; backtest_paired_metrics: {_n_pairs} pairs (written)")
else:
print(
f"cohort_metrics: {_n_cohorts} rows; backtest_paired_metrics: {_n_pairs} pairs "
"(already present, nothing written)"
)
```
## §1 Handoff from model analysis
The strategy phase inherits one selected model from `10_model_analysis.py`
§8. That selection is what §3's deflated Sharpe and §6's holdout closure are
correcting for, so it is stated here rather than treated as settled.
The rule that makes the selection is the pipeline's, not this notebook's: the
highest validation backtest Sharpe across the baseline and allocation stages.
Nothing here overrides it, so which family, configuration and label come
through is read off the resolver below and printed rather than asserted in
prose. IC on the prediction set is computed on continuous model scores against
continuous realized returns.
All three declared labels are backtested: the setup-primary `fwd_ret_1m`, the
winsorized `fwd_ret_1m_win` and the classification variant `fwd_class_1m`. They
compete at the baseline and the selection ranges over all of them, so the label
the strategy phase carries is read off the resolver below rather than assumed
to be the primary one.
The shared resolver applies the common-support contract whenever corrected
conformal rows are present: those candidates start one fold later than the
equal-weight baselines, so ranking them as registered would credit the
difference to the sizing rule when part of it is the shorter window. Validation
chooses the strategy once, and the holdout takes no part in that ranking.
Which sizing rule comes through is read from the selected run's own strategy
spec rather than assumed, because both the equal-weight baseline and the
allocator variants are candidates and either can win.
```python
rank1 = resolve_solvent_carrier(CASE_STUDY)
TOP_HASH = rank1["val_backtest_hash"]
TOP_PHASH = rank1["val_prediction_hash"]
RANK1_STAGE = rank1["val_stage"]
RANK1_FAMILY = rank1["family"]
RANK1_CONFIG = rank1["config_name"]
RANK1_LABEL = rank1["label"]
COMMON_SUPPORT_SHARPE = rank1["val_sharpe"]
COMMON_SUPPORT_N = rank1["comparison_n_periods"]
COMMON_SUPPORT_START = rank1["comparison_start"]
COMMON_SUPPORT_END = rank1["comparison_end"]
with sqlite3.connect(str(CASE_DIR / "run_log" / "registry.db")) as _con:
_spec = json.loads(
_con.execute(
"SELECT spec_json FROM backtest_runs WHERE backtest_hash = ?", (TOP_HASH,)
).fetchone()[0]
)
RANK1_ALLOCATOR = (_spec.get("strategy", {}).get("allocation") or {}).get("method", "equal_weight")
RANK1_TOP_K = int(_spec["strategy"]["signal"]["top_k"])
```
Registry metadata anchors the selected run and its prediction-side IC to the
exact validation hashes used downstream.
```python
with sqlite3.connect(str(_db)) as _con:
_row = _con.execute(
"SELECT ic_mean_daily, ic_ci_lo, ic_ci_hi, ic_t_hac, ic_p_hac, ic_n_days, "
"ic_hac_lag, ic_pct_positive "
"FROM prediction_metrics WHERE prediction_hash = ?",
(TOP_PHASH,),
).fetchone()
```
Selection-bias statistics come from the family cohort that contains the
selected validation run.
```python
with sqlite3.connect(str(_db)) as _con:
_cohort_row = _con.execute(
"SELECT k_variants, dsr_raw, dsr_raw_pvalue, dsr_mp, dsr_mp_pvalue, "
"dsr_er, dsr_er_pvalue, expected_max_sharpe_raw, "
"expected_max_sharpe_mp, expected_max_sharpe_er, min_trl_periods_raw, "
"min_trl_periods_mp, min_trl_periods_er, pbo, pbo_n_combinations, "
"pbo_n_folds, leader_hash, leader_sharpe "
"FROM cohort_metrics "
"WHERE cohort_type='family' AND stage=? AND label=? AND family=?",
(RANK1_STAGE, RANK1_LABEL, RANK1_FAMILY),
).fetchone()
```
```python
COHORT = (
{
"k_variants": _cohort_row[0],
"dsr_raw": _cohort_row[1],
"dsr_raw_pvalue": _cohort_row[2],
"dsr_mp": _cohort_row[3],
"dsr_mp_pvalue": _cohort_row[4],
"dsr_er": _cohort_row[5],
"dsr_er_pvalue": _cohort_row[6],
"expected_max_sharpe_raw": _cohort_row[7],
"expected_max_sharpe_mp": _cohort_row[8],
"expected_max_sharpe_er": _cohort_row[9],
"min_trl_periods_raw": _cohort_row[10],
"min_trl_periods_mp": _cohort_row[11],
"min_trl_periods_er": _cohort_row[12],
"pbo": _cohort_row[13],
"pbo_n_combinations": _cohort_row[14],
"pbo_n_folds": _cohort_row[15],
"leader_hash": _cohort_row[16],
"leader_sharpe": _cohort_row[17],
}
if _cohort_row is not None
else None
)
```
```python
ic_mean, ic_lo, ic_hi, ic_t, ic_p, ic_ndays, ic_lag, ic_pct = _row
print(f"Rank-1: family={RANK1_FAMILY}, config={RANK1_CONFIG}, label={RANK1_LABEL}")
print(f" prediction_hash={TOP_PHASH}, backtest_hash={TOP_HASH}")
print(
f" common-support Sharpe={COMMON_SUPPORT_SHARPE:.4f}, "
f"n={COMMON_SUPPORT_N}, {COMMON_SUPPORT_START:%Y-%m-%d} to {COMMON_SUPPORT_END:%Y-%m-%d}"
)
print(
f" setup-primary label={PRIMARY_LABEL}; comparisons follow rank-1 label={RANK1_LABEL}."
)
print()
print("Per-decision-time IC (validation, continuous score vs continuous returns):")
print(f" IC = {_fmt_ci(ic_mean, ic_lo, ic_hi, '.4f')} (HAC, lag={int(ic_lag)})")
print(f" t_HAC = {ic_t:.3f}, p_HAC = {ic_p:.3e}")
print(f" n_days = {int(ic_ndays)}, pct_positive = {ic_pct:.1%}")
print(f" CI status: {ci_status(ic_lo, ic_hi)}")
```
The holdout run is determined by the selected validation run rather than chosen. The
resolver requires an exact strategy-specification match, the same checkpoint, and a
training run whose own CV declares the holdout fold - so what it returns is
[`15_holdout_predictions`](15_holdout_predictions.ipynb)' refit traded by
[`16_holdout_backtest`](16_holdout_backtest.ipynb), and not a model fitted on the
validation folds and scored here. Any looser rule would let an allocator be picked on
its holdout performance, which is selection wearing the clothes of evaluation.
```python
HO_HASH = select_holdout_self_backtest(CASE_STUDY, TOP_HASH)
if HO_HASH is None:
raise RuntimeError("No strict-lineage holdout matches the selected validation run.")
if rank1["holdout_backtest_hash"] != HO_HASH:
raise RuntimeError("Canonical resolver and exact-spec holdout lookup disagree.")
with sqlite3.connect(str(_db)) as _con:
_holdout_meta = _con.execute(
"SELECT p.training_hash, t.family, t.config_name, t.label, pm.ic_mean_daily "
"FROM backtest_runs b "
"JOIN prediction_sets p ON b.prediction_hash=p.prediction_hash "
"JOIN training_runs t ON p.training_hash=t.training_hash "
"JOIN prediction_metrics pm ON p.prediction_hash=pm.prediction_hash "
"WHERE b.backtest_hash=? AND p.split='holdout'",
(HO_HASH,),
).fetchone()
_HO_TRAINING_HASH, HO_FAMILY, HO_CONFIG, HO_LABEL, HO_IC = _holdout_meta
```
The publication registry is complete before this analysis runs. The two
derived holdout comparisons must already exist and the validation comparison
must name the frozen validation run exactly.
```python
with sqlite3.connect(str(_db)) as _con:
_holdout_pairs = _con.execute(
"SELECT benchmark_kind, benchmark_hash FROM backtest_paired_metrics "
"WHERE challenger_hash=? AND benchmark_kind IN "
"('val_rank1_self', 'equal_weight_holdout_side_artifact')",
(HO_HASH,),
).fetchall()
_pair_map = {kind: benchmark_hash for kind, benchmark_hash in _holdout_pairs}
if set(_pair_map) != {"val_rank1_self", "equal_weight_holdout_side_artifact"}:
raise RuntimeError(f"Incomplete holdout closure pairs: {_pair_map}")
if _pair_map["val_rank1_self"] != TOP_HASH:
raise RuntimeError("Holdout decay is not paired against the frozen validation run.")
```
```python
print(
f"Strict-lineage holdout: {HO_FAMILY}/{HO_CONFIG}, label={HO_LABEL}, "
f"backtest_hash={HO_HASH}, IC={HO_IC:.4f}"
)
```
The IC table above reads as a continuous-score-vs-continuous-return
signal: HAC-adjusted CI status indicates whether the prediction edge
is `excludes_zero_strong` on the positive side. The strategy-side
translation of the prediction-stage IC depends on how concentrated the
book is: the same IC spread over a handful of names and over the whole
cross section produce different Sharpes. §3 reports the selection-adjusted
Sharpe under the sweep's k-variants search.
**Kill conditions** are not encoded in `setup.yaml` for this CS. The
notebook evaluates two universal gates in §9: (i) the validation Sharpe
CI lower bound ≥ 0, and (ii) the holdout strategy-vs-EW paired CI does
not exclude zero on the negative side. Both are reported as
pass, partial or fail, and neither is turned into a recommendation.
## §2 Search context, family comparison, and lineage waterfall
The equal-weight baseline covers the full family x config x label grid. The
search-context table below summarizes that distribution; the family-level
breakdown then locates the selected run within it, which is what makes the size
of the search visible rather than only its maximum.
```python
ctx = explorer.search_context("signal")
search_table = pl.DataFrame(
[
{"metric": "Total baseline backtests", "value": f"{ctx['total']:,}"},
{"metric": "Mean Sharpe", "value": f"{ctx['mean_sharpe']:.3f}"},
{"metric": "Median Sharpe", "value": f"{ctx['median_sharpe']:.3f}"},
{"metric": "P90 Sharpe", "value": f"{ctx['p90_sharpe']:.3f}"},
{"metric": "% positive Sharpe", "value": f"{ctx['pct_positive']:.1f}%"},
{"metric": "Top-by-Sharpe in this sweep", "value": f"{ctx['champion_sharpe']:.3f}"},
{"metric": "Top-by-Sharpe percentile", "value": f"{ctx['champion_percentile']:.1f}%"},
]
)
print("Equal-weight baseline search context:")
print(search_table)
```
```python
with sqlite3.connect(str(_db)) as _con:
_famdf = pl.DataFrame(
_con.execute(
"""
SELECT
t.family,
bm.sharpe,
bm.max_drawdown,
bm.sharpe_ci95_lo,
bm.sharpe_ci95_hi
FROM backtest_metrics bm
JOIN backtest_runs b ON bm.backtest_hash = b.backtest_hash
JOIN prediction_sets p ON b.prediction_hash = p.prediction_hash
JOIN training_runs t ON p.training_hash = t.training_hash
WHERE b.stage = 'signal'
AND p.split = 'validation'
AND (bm.num_trades IS NULL OR bm.num_trades > 0)
"""
).fetchall(),
schema=["family", "sharpe", "max_drawdown", "sharpe_ci95_lo", "sharpe_ci95_hi"],
orient="row",
)
# The statistics are taken over the runs that stayed solvent, and the rest are counted
# beside them - the same treatment `11_backtest` gives its family table, so the two agree
# rather than reporting different medians for the same family. A run whose equity reached
# zero has no capital left to earn a return, so its Sharpe is arithmetic on nothing and
# would move a median it has no claim on; dropping those runs without counting them would
# instead rank a family by its survivors. Runs with no recorded drawdown are counted apart,
# because a failure that was never measured is not a bankruptcy. The rows are not filtered
# on Sharpe: a bankrupt run can carry a null one, and dropping it in SQL would remove it
# before `insolvent` counts it.
_solvent = pl.col("max_drawdown") > -1.0
_solvent_mask = _solvent.fill_null(False)
family_summary = (
_famdf.group_by("family")
.agg(
n=_solvent.sum(),
insolvent=(pl.col("max_drawdown") <= -1.0).sum(),
unknown=pl.col("max_drawdown").is_null().sum(),
sharpe_median=pl.col("sharpe").filter(_solvent_mask).median(),
sharpe_q25=pl.col("sharpe").filter(_solvent_mask).quantile(0.25),
sharpe_q75=pl.col("sharpe").filter(_solvent_mask).quantile(0.75),
sharpe_max=pl.col("sharpe").filter(_solvent_mask).max(),
pct_positive=(pl.col("sharpe").filter(_solvent_mask) > 0).mean() * 100,
)
.sort("sharpe_median", descending=True, nulls_last=True)
)
print("Family-level equal-weight baseline Sharpe summary, over the solvent runs:")
# Nine columns is past the default width, and the column polars elides first is the median -
# the one the text above asks the reader to compare against the maximum.
with pl.Config(tbl_cols=-1, tbl_width_chars=160):
print(family_summary)
```
```python
fig, ax = plt.subplots(figsize=(9, 4))
fams = family_summary["family"].to_list()
y = np.arange(len(fams))
medians = family_summary["sharpe_median"].to_numpy()
q25 = family_summary["sharpe_q25"].to_numpy()
q75 = family_summary["sharpe_q75"].to_numpy()
maxima = family_summary["sharpe_max"].to_numpy()
ax.errorbar(
medians,
y,
xerr=[medians - q25, q75 - medians],
fmt="o",
color=COLORS["blue"],
ecolor=COLORS["slate"],
elinewidth=2.0,
capsize=4,
label="median with IQR",
)
ax.scatter(maxima, y, marker="x", color=COLORS["amber"], s=60, label="max", zorder=5)
ax.axvline(0, color=COLORS["neutral"], linewidth=0.8, linestyle="--")
ax.set_yticks(y)
ax.set_yticklabels(fams)
ax.set_xlabel("Validation Sharpe")
add_message_title(ax, "One family carries the upper tail of the baseline sweep")
ax.invert_yaxis()
ax.legend(loc="lower right", frameon=False)
show_with_alt(
fig,
"Validation Sharpe by model family, one row per family: a marker at the median with a bar "
"spanning the interquartile range, a cross at the family maximum, and a dashed line at zero.",
)
```
Two statistics answer different questions here. The median says whether a
family's configurations were generally worth trading; the maximum says only
how far its luckiest draw reached, and it rises with the number of
configurations tried whether or not the family has an edge. Where the
families' interquartile ranges overlap, the selected run's margin is a matter of
where it sat within its own family's spread rather than of which family it
belongs to.
The lineage records all stages registered against the selected prediction set.
The stage named `signal` in the registry is the equal-weight baseline.
```python
lineage = explorer.champion_lineage(TOP_PHASH)
ci_lo: dict[str, float] = {}
ci_hi: dict[str, float] = {}
for stage_name, info in lineage.items():
bm = load_backtest_metrics(CASE_STUDY, backtest_hash=info["backtest_hash"])
if not bm.is_empty():
row = bm.row(0, named=True)
ci_lo[stage_name] = row.get("sharpe_ci95_lo")
ci_hi[stage_name] = row.get("sharpe_ci95_hi")
print("Lineage stages present for the selected prediction:")
for s, info in lineage.items():
lo, hi = ci_lo.get(s), ci_hi.get(s)
print(f" {s}: hash={info['backtest_hash']}, Sharpe={_fmt_ci(info['sharpe'], lo, hi)}")
missing_stages = [s for s in ("allocation", "cost_sensitivity", "risk_overlay") if s not in lineage]
if missing_stages:
print()
print(f"Stage transitions not run for this prediction: {missing_stages}")
print(
" (cost_sensitivity and risk_overlay sweeps exist in the registry but "
"were registered against different prediction_hashes; cross-prediction "
"stage transitions are Ch20's lane.)"
)
```
The corrected conformal candidates start one fold later than the equal-weight
baselines, so their registry windows are not the same length. Comparing the two
Sharpes as registered would credit the difference to the sizing rule when part
of it is the window. The comparison below recomputes both on the exact
timestamp intersection and raises if the two supports still differ.
One row per sizing rule the eligible population registers, each the strongest
candidate under that rule, found by ranking rather than named - so the comparison
reads the same way whichever rule the selection landed on, and adding an allocator
to the sweep adds a row here rather than silently falling outside the check.
They are drawn from the population the selection rule itself draws from, and they have
to be: both arms are compared against the selected run, and the check below requires
that run to be the strongest under its own sizing rule. A wider pool fails that check
on a run that was never a candidate.
The cost-sensitivity stage is what has to be excluded. Those rows are this same
strategy re-run at each level of the cost grid, so the zero-cost row scores above its
own declared-cost parent, and ranking them here would treat a cost assumption as a
sizing rule. Benchmark families and prediction sets carrying a constant-prediction fold are
excluded for the reason the resolver excludes them: neither is eligible to carry the
case study.
```python
_ELIGIBLE_STAGES = ("signal", "allocation", "risk_overlay", "holdout")
with sqlite3.connect(str(_db)) as _con:
_sizing_candidates = _con.execute(
"SELECT b.backtest_hash, "
" COALESCE(json_extract(b.spec_json, '$.strategy.allocation.method'),"
" 'equal_weight') AS allocator, "
" json_extract(b.spec_json, '$.strategy.allocation.calibration_version') AS calibration "
"FROM backtest_runs b "
"JOIN backtest_metrics bm ON bm.backtest_hash=b.backtest_hash "
"JOIN prediction_sets p ON p.prediction_hash=b.prediction_hash "
"JOIN training_runs t ON t.training_hash=p.training_hash "
f"WHERE b.stage IN ({', '.join('?' * len(_ELIGIBLE_STAGES))}) "
"AND p.split='validation' AND bm.sharpe IS NOT NULL "
"AND t.family != 'benchmark' "
# The same live population the cohort corrections are scoped to. Unscoped, this pool
# holds superseded generations the selection never ranked, and they decide how far the
# common-support intersection reaches - so the comparison would be computed over a
# window the selected run was never ranked on, and the chart would label it with the
# resolver's period count while showing another.
f"AND p.prediction_hash IN ({', '.join('?' * len(LIVE_PREDICTIONS))})"
+ degenerate_prediction_sql("p.prediction_hash"),
(*_ELIGIBLE_STAGES, *LIVE_PREDICTIONS),
).fetchall()
# A superseded conformal calibration is not a candidate; no other sizing rule has such a
# distinction to make. The current contract is read from `CALIBRATION_VERSION` rather than
# written out here: this line named `walk_forward_v2` as the strict rule, and when the
# fleet moved to `walk_forward_v3` the literal stopped matching anything and the notebook
# refused with "no conformal candidate is registered" - which reads as an absent sweep and
# was a stale string.
#
# Which sizing rules exist is read from the registry rather than named here. This cell named
# `equal_weight` and `conformal_weighted` and built a comparison row for each, while the
# selection rule ranges over the whole allocator sweep. When the winsorized label's
# `score_weighted` allocation took rank 1, the selected run belonged to a rule the comparison
# had no row for, and the check below refused it for disagreeing with a ranking it was never
# entered in. Reading the rules off the same population the selection draws from is what makes
# "strongest under its own sizing rule" a statement about the selection rather than about
# which two rules this cell happened to list.
_method_of = {
row[0]: row[1]
for row in _sizing_candidates
if row[1] != "conformal_weighted" or row[2] == CALIBRATION_VERSION
}
_methods = set(_method_of.values())
if "conformal_weighted" not in _methods:
raise RuntimeError(
f"No conformal candidate at the current calibration {CALIBRATION_VERSION!r} is "
"registered. Re-run 12_portfolio_management: the registry's conformal generation "
"predates the current contract and cannot be executed."
)
if "equal_weight" not in _methods:
raise RuntimeError("No equal-weight baseline candidate is registered.")
if TOP_HASH not in _method_of:
raise RuntimeError(
f"The selected run {TOP_HASH} is not in the eligible sizing population, so this "
"comparison would rank it against candidates the selection never saw."
)
sizing_ranking = rank_backtests_on_common_support(
CASE_STUDY,
sorted(_method_of),
periods_per_year=PERIODS_PER_YEAR,
)
sizing_ranking = sizing_ranking.with_columns(
pl.col("backtest_hash").replace_strict(_method_of).alias("method")
)
selection_comparison = (
sizing_ranking.sort("sharpe", descending=True)
.group_by("method", maintain_order=True)
.first()
.with_columns(
pl.when(pl.col("backtest_hash") == TOP_HASH)
.then(pl.col("method") + pl.lit(" (selected)"))
.otherwise(pl.col("method"))
.alias("method")
)
)
if selection_comparison["n_periods"].n_unique() != 1:
raise RuntimeError("Selection comparison does not have identical period support.")
if TOP_HASH not in selection_comparison["backtest_hash"].to_list():
raise RuntimeError(
"The selected run is not the strongest under its own sizing rule on common support, "
"which means the registered ranking and the common-support ranking disagree."
)
print("Strongest candidate under each sizing rule, on common support:")
print(selection_comparison.select("method", "backtest_hash", "sharpe", "n_periods", "start", "end"))
```
```python
stage_labels = selection_comparison["method"].to_list()
stage_values = selection_comparison["sharpe"].to_numpy()
fig, ax = plt.subplots(figsize=(8, 4.5), constrained_layout=True)
x = np.arange(len(stage_labels))
bars = ax.bar(
x,
stage_values,
color=[COLORS["blue"], COLORS["amber"]],
width=0.55,
)
ax.bar_label(bars, labels=[f"{value:.3f}" for value in stage_values], padding=4)
ax.axhline(0, color=COLORS["neutral"], linewidth=0.8)
ax.set_xticks(x, stage_labels)
_COMPARISON_N = int(selection_comparison["n_periods"][0])
ax.set_ylabel(f"Validation Sharpe on common {_COMPARISON_N}-month support")
add_message_title(
ax,
"Selection reads the peak of each sizing rule, not its average",
subtitle="Strongest candidate under each rule, recomputed on the shared window",
)
show_with_alt(
fig,
"One bar per sizing rule, each labelled with the Sharpe of the strongest candidate under that "
"rule, recomputed on the months the two rules share.",
)
```
Two things are true of this comparison and the second does not retract the
first.
The selection rule is the highest validation backtest Sharpe, and it is applied
here without an override. Whichever rule holds the higher bar above is the one
the strategy phase carries, and §3 onward describe that run.
The rule reads one number per candidate, the highest it reached. It does not read
how that rule did across the rest of the grid, and `12_portfolio_management`
measured exactly that. Averaged over the same ten predictions at the same four
concentrations, conformal weighting came out below equal weighting, and eight
of its ten runs at the narrowest concentration ended with negative equity
against equal weighting's two. A rule can own the single highest cell and still
be the weaker rule across the grid, and a selection made on the maximum cannot
tell the two apart.
That is a property of selecting on a maximum rather than a defect in this
case study, and it is the reason §3 reports a deflated Sharpe and §6 closes on
a holdout the selection never saw. Neither correction recovers the average the
selection did not look at; they bound how much of the maximum was the search.
§5 reads the cost-sensitivity sweep as a friction floor across the allocation
cohort rather than as a transition on this lineage.
## §3 Headline performance with uncertainty
The specification is chosen on the common comparison window, and the metric
table then reports the selected run's complete validation record with
block-bootstrap confidence intervals from `backtest_metrics`. Those two windows are not
the same length; the spec block above prints both so the reader can see by
how much.
The bootstrap block length captures serial dependence in the return series and is not
required to match the one-month rebalance step, but it cannot be shorter than it: a
block inside a holding period resamples within one position and understates the
dependence the interval exists to carry. The audit at the end of the cell below
computes that relation rather than asserting it, and stops the notebook if it fails.
```python
full = load_backtest_metrics(CASE_STUDY, backtest_hash=TOP_HASH).row(0, named=True)
top_strategy = _spec["strategy"]
top_signal = top_strategy["signal"]
top_allocation = top_strategy.get("allocation") or {}
spec_block = {
"case_study": CASE_STUDY,
"family": RANK1_FAMILY,
"config_name": RANK1_CONFIG,
"label": RANK1_LABEL,
"setup_primary_label": PRIMARY_LABEL,
"signal_method": top_signal.get("method"),
"allocator": top_allocation.get("method", "equal_weight"),
"top_k": top_signal.get("top_k"),
"rebalance_step": setup["labels"]["rebalance_step"][RANK1_LABEL],
"cost_assumption": "stage-dependent; §5 reports the cost sensitivity curve",
"cross_stage_rank1_stage": RANK1_STAGE,
"selection_sharpe_common_support": COMMON_SUPPORT_SHARPE,
"selection_window_periods": COMMON_SUPPORT_N,
"full_validation_sharpe": full["sharpe"],
"full_validation_periods": int(full["n_periods"]),
"num_trades": int(full["num_trades"]) if full["num_trades"] is not None else None,
"avg_turnover": full["avg_turnover"],
"bootstrap_block_length": int(full["bootstrap_block_length"]),
"bootstrap_n": int(full["bootstrap_n"]),
}
print("Selected run, validation window:")
for k, v in spec_block.items():
print(f" {k}: {v}")
_block = int(full["bootstrap_block_length"])
_rstep = setup["labels"]["rebalance_step"][RANK1_LABEL]
if _block < _rstep:
raise RuntimeError(
f"bootstrap block is {_block} months against a {_rstep}-month rebalance step. A "
"block shorter than the holding period resamples inside it, so the confidence "
"intervals below would understate the serial dependence they exist to carry."
)
print(
f" dependence audit: bootstrap block={_block} months, rebalance step={_rstep} month, "
"so the block spans at least one holding period"
)
```
A compact row helper keeps the metric table construction readable.
```python
def _row(metric: str, point: str, lo: str, hi: str, status: str) -> dict:
return {"metric": metric, "point": point, "ci95_lo": lo, "ci95_hi": hi, "status": status}
```
```python
metric_specs = [
("Sharpe", "sharpe", "sharpe_ci95_lo", "sharpe_ci95_hi"),
("Sortino", "sortino", "sortino_ci95_lo", "sortino_ci95_hi"),
("Annualized return", "cagr", "ann_return_ci95_lo", "ann_return_ci95_hi"),
("Max drawdown", "max_drawdown", "max_dd_ci95_lo", "max_dd_ci95_hi"),
("Calmar", "calmar", "calmar_ci95_lo", "calmar_ci95_hi"),
]
headline_rows = [
_row(name, _fmt(full[point]), _fmt(full[lo]), _fmt(full[hi]), ci_status(full[lo], full[hi]))
for name, point, lo, hi in metric_specs
]
headline_rows.append(_row("PSR p-value (H0: SR <= 0)", _fmt(full["psr_pvalue"]), "-", "-", "n/a"))
for name, key, p_key, prefix in [
("DSR raw", "dsr_raw", "dsr_raw_pvalue", "raw"),
("DSR ER (maintainer default)", "dsr_er", "dsr_er_pvalue", "er"),
("DSR MP", "dsr_mp", "dsr_mp_pvalue", "mp"),
]:
point = _fmt(COHORT[key]) if COHORT else "unavailable"
status = (
f"{prefix}_p={COHORT[p_key]:.3f}" if COHORT and COHORT[p_key] is not None else "unavailable"
)
headline_rows.append(_row(name, point, "-", "-", status))
headline_rows.extend(
[
_row("Expected max Sharpe (ER)", _fmt(COHORT["expected_max_sharpe_er"]), "-", "-", "n/a"),
_row("PBO", _fmt(COHORT["pbo"]), "-", "-", "n/a"),
_row("k variants (family cohort)", str(int(COHORT["k_variants"])), "-", "-", "n/a"),
]
)
headline = pl.DataFrame(headline_rows)
print("Selected strategy, headline metrics with confidence intervals:")
print(headline)
```
Each interval below is drawn from its own lower bound to its own upper bound, with the
point estimate marked on it, rather than as an error bar measured outwards from the
point. The two look identical whenever the interval brackets its estimate and differ
only when it does not - and drawing it this way means that case renders and is visible
instead of raising. A percentile bootstrap is not guaranteed to contain the full-sample
estimate on a skewed series, so an estimate outside its interval is worth reporting and
is not on its own proof of a defect. It was proof of one here: the stored Sortino ratio
and the bootstrap around it were computed from two different downside deviations, and
the printed line below says which metrics sit outside their own bounds.
```python
ew_val = load_benchmark_metrics(CASE_STUDY, RANK1_LABEL, period="validation")
forest_metrics = [
("Sharpe", full["sharpe"], full["sharpe_ci95_lo"], full["sharpe_ci95_hi"]),
("Sortino", full["sortino"], full["sortino_ci95_lo"], full["sortino_ci95_hi"]),
("Calmar", full["calmar"], full["calmar_ci95_lo"], full["calmar_ci95_hi"]),
]
inverted = [
f"{name} ({point:.4f} against [{lo:.4f}, {hi:.4f}])"
for name, point, lo, hi in forest_metrics
if not (lo <= point <= hi)
]
print(
"Every estimate lies inside its own interval."
if not inverted
else "Outside its own interval, so the estimate and the interval may not be the same "
f"quantity measured two ways: {'; '.join(inverted)}"
)
fig, ax = plt.subplots(figsize=(8, 4))
y = np.arange(len(forest_metrics))
points = np.array([m[1] for m in forest_metrics])
los = np.array([m[2] for m in forest_metrics])
his = np.array([m[3] for m in forest_metrics])
ax.hlines(y, los, his, color=COLORS["slate"], linewidth=2.0)
ax.scatter(points, y, color=COLORS["blue"], s=49, zorder=3)
ax.axvline(0, color=COLORS["neutral"], linestyle="--", linewidth=0.8)
ax.set_yticks(y)
ax.set_yticklabels([m[0] for m in forest_metrics])
ax.invert_yaxis()
ax.set_xlabel("Risk-adjusted ratio")
_above_zero = sum(1 for _, _, lo, _ in forest_metrics if lo > 0)
add_message_title(
ax,
f"{_above_zero} of {len(forest_metrics)} risk-adjusted metrics hold a lower bound above zero"
if _above_zero < len(forest_metrics)
else "Every risk-adjusted metric holds a lower bound above zero",
)
show_with_alt(
fig,
"One row per risk-adjusted metric, each a point estimate on a horizontal line spanning its "
"interval, against a dashed line at zero.",
)
```
```python
strat_returns_path = CASE_DIR / "run_log" / "backtest" / TOP_HASH / "daily_returns.parquet"
strat_df = (
pl.read_parquet(strat_returns_path)
.sort("timestamp")
.with_columns(pl.col("timestamp").cast(pl.Date).alias("ts"))
.select(pl.col("ts"), pl.col("daily_return").alias("strategy"))
)
bench_val = (
load_benchmark_returns(CASE_STUDY, RANK1_LABEL, period="validation")
.with_columns(pl.col("timestamp").cast(pl.Date).alias("ts"))
.select(pl.col("ts"), pl.col("ew_return").alias("benchmark"))
)
aligned = strat_df.join(bench_val, on="ts", how="inner").sort("ts")
print(
f"Validation overlay window: {aligned['ts'].min()} → {aligned['ts'].max()}, n={aligned.height}"
)
```
The validation Sharpe CI is reported above with its `ci_status`
classification. PSR p-value tests H0: Sharpe ≤ 0; the DSR variants
above (raw / MP / ER) report the selection-adjusted equivalent across
the k_variants in the family cohort, with ER the
maintainer-recommended default. PBO is the probability of backtest
overfitting from CSCV on the per-fold Sharpe matrix; §6's val→ho
closure provides the orthogonal out-of-sample check.
The cumulative-return plot shows the validation trajectory against the EW
universe. The y-axis is the cumulative product of monthly returns from the
dollar-neutral long-short structure, so across a decade of months it reaches
end-points that are arithmetic rather than attainable: they assume every
month's gain is reinvested at the same leverage with no capacity limit and
no borrow constraint. The headline annualized-return CI inherits that
assumption. §5 puts the gross profile against implementation friction.
### Second benchmark: Fama-French market
The EW universe is the inside-the-strategy benchmark (untimed
allocation across the same names). The Fama-French market series
(Mkt-RF + RF) is the classical academic external benchmark - what
every cross-sectional anomaly is implicitly compared against. It
matches the strategy's monthly cadence directly. Beta against the
market quantifies whether the long-short construction achieves the
market-neutrality it nominally targets.
```python
ff_m = load_ff_market_returns(
start=aligned["ts"].min(),
end=aligned["ts"].max(),
frequency="monthly",
)
ff_aligned = align_external_benchmark_monthly(
aligned.select("ts", "strategy"),
ff_m,
timestamp_col="ts",
)
ff_diag = compute_benchmark_diagnostics(
ff_aligned["strategy"].to_numpy(),
ff_aligned["benchmark_return"].to_numpy(),
PERIODS_PER_YEAR,
)
ew_diag = compute_benchmark_diagnostics(
aligned["strategy"].to_numpy(),
aligned["benchmark"].to_numpy(),
PERIODS_PER_YEAR,
)
```
```python
benchmark_table = pl.DataFrame(
[
{
"benchmark": "EW universe (validation)",
"n_periods": ew_diag["n"],
"info_ratio": _fmt(ew_diag["info_ratio"], ".3f"),
"beta": _fmt(ew_diag["beta"], ".3f"),
"correlation": _fmt(ew_diag["correlation"], ".3f"),
"tracking_error": _fmt(ew_diag["tracking_error"], ".3f"),
},
{
"benchmark": "FF-market (Mkt-RF + RF, monthly)",
"n_periods": ff_diag["n"],
"info_ratio": _fmt(ff_diag["info_ratio"], ".3f"),
"beta": _fmt(ff_diag["beta"], ".3f"),
"correlation": _fmt(ff_diag["correlation"], ".3f"),
"tracking_error": _fmt(ff_diag["tracking_error"], ".3f"),
},
]
)
print("Strategy diagnostics vs benchmarks:")
print(benchmark_table)
```
### Sub-period decomposition (5-year buckets)
A pooled Sharpe averages over the whole validation window, and an average
hides whether the return arrived steadily or in one stretch. The buckets
below split validation into two roughly five-year windows and place the
holdout beside them. The holdout is resolved through the unique
complete-strategy-spec match within the frozen training lineage, the same
rule §6 uses for the paired closure.
```python
ho_panel = (
pl.read_parquet(CASE_DIR / "run_log" / "backtest" / HO_HASH / "daily_returns.parquet")
.sort("timestamp")
.with_columns(pl.col("timestamp").cast(pl.Date).alias("ts"))
.select("ts", pl.col("daily_return").alias("strategy"))
)
bench_ho = (
load_benchmark_returns(CASE_STUDY, RANK1_LABEL, period="holdout")
.with_columns(pl.col("timestamp").cast(pl.Date).alias("ts"))
.select("ts", pl.col("ew_return").alias("benchmark"))
)
ho_aligned = (
ho_panel.join(bench_ho, on="ts", how="inner").sort("ts")
if ho_panel.height > 0
else pl.DataFrame()
)
val_buckets = [
("2006-2010 (val)", _dt.date(2006, 1, 1), _dt.date(2010, 12, 31)),
("2011-2015 (val)", _dt.date(2011, 1, 1), _dt.date(2015, 12, 31)),
]
val_table = compute_subperiod_diagnostics(
aligned,
val_buckets,
periods_per_year=PERIODS_PER_YEAR,
)
ho_table = (
compute_subperiod_diagnostics(
ho_aligned,
[("2016 (ho)", _dt.date(2016, 1, 1), _dt.date(2016, 12, 31))],
periods_per_year=PERIODS_PER_YEAR,
)
if ho_aligned.height > 0
else pl.DataFrame()
)
subperiod_table = pl.concat([val_table, ho_table]) if ho_table.height > 0 else val_table
print("Sub-period decomposition (validation 5y buckets + holdout):")
print(subperiod_table)
```
## §4 Risk and drawdown analysis
Risk metrics use the validation-window strategy returns paired against
the validation EW benchmark. The headline CIs and tail-risk read
cover the substantive risk profile; per-fold decomposition (where
available in `backtest_fold_metrics`) localizes the realized Sharpe
across the CV folds.
```python
strat_arr = aligned["strategy"].to_numpy()
bench_arr = aligned["benchmark"].to_numpy()
ts_arr = aligned["ts"].to_list()
pa = PortfolioAnalysis(
returns=strat_arr,
benchmark=bench_arr,
dates=ts_arr,
periods_per_year=PERIODS_PER_YEAR,
)
dd = pa.compute_drawdown_analysis()
print("Drawdown analysis (validation window):")
print(dd)
```
```python
tail_table = pl.DataFrame(
[
{"metric": "Volatility (ann.)", "value": f"{full['volatility']:.4f}"},
{"metric": "VaR 95% (period)", "value": f"{full['var_95']:.4f}"},
{"metric": "CVaR 95% (period)", "value": f"{full['cvar_95']:.4f}"},
{"metric": "Tail ratio", "value": f"{full['tail_ratio']:.3f}"},
{"metric": "Skewness", "value": f"{full['skewness']:.3f}"},
{"metric": "Kurtosis", "value": f"{full['kurtosis']:.3f}"},
]
)
print()
print("Tail risk profile:")
print(tail_table)
```
```python
fold_df = load_backtest_fold_metrics(CASE_STUDY, backtest_hash=TOP_HASH)
print(f"Per-fold breakdown: {fold_df.height} rows.")
if fold_df.is_empty():
print(
" (Empty - fold writer did not register fold metrics for this "
"lineage. Headline CI in §3 stands in for fold variance.)"
)
else:
print(fold_df.select("fold_id", "sharpe", "max_drawdown", "n_days"))
```
Max drawdown is the worst peak-to-trough excursion the strategy actually
experienced, and its lower confidence bound is the one to plan against: it
is the depth a run of this length could plausibly have produced, so a
deployment sized to survive the point estimate but not the lower bound is
undercapitalized. Annualized volatility is stated under the dollar-neutral
leverage assumption; tail kurtosis and the tail ratio describe how
symmetric the realized distribution was.
## §4b Inline diagnostic panels
§8 generates the standalone tear sheet HTML when feasible (the
Plotly Dash design renders best as a multi-tab app). For inline
review we lift the diagnostic library's core panels off the
validation `PortfolioAnalysis`. Trade-conditional panels - IC time
series, decile returns, prediction–trade alignment - are
**unavailable** for this case study: the vectorized backtester used
here emits `daily_returns.parquet` and `weights.parquet` but no
`trades.parquet`, and the bridge layer that builds a `BacktestProfile`
requires per-trade fills to construct the ML surfaces. Per the strategy-analysis
convention we do not synthesize fills. The IC analysis covered in
§3 (forest plot) and §6 (paired holdout) is the substitute; the
return-distribution and risk views below complete the inline picture.
**Cumulative return.** Validation-window equity vs. EW universe.
```python
fig_cum = plot_cumulative_returns(pa, benchmark_label="EW universe")
fig_cum.update_layout(title="The selected strategy compounds ahead of the universe benchmark")
fig_cum.show()
```
**Annual returns.** Calendar-year strategy vs. benchmark bars.
```python
fig_annual = plot_annual_returns_bar(pa, benchmark_label="EW universe")
fig_annual.update_layout(title="The strategy posts gains in every observed validation year")
fig_annual.show()
```
**Monthly returns heatmap.** Year×month return surface.
```python
fig_monthly = plot_monthly_returns_heatmap(pa)
fig_monthly.update_layout(title="Gains persist across most validation months")
fig_monthly.show()
```
**Returns distribution.** Monthly return histogram against a normal
overlay; complements §4's tail-ratio table.
```python
fig_dist = plot_returns_distribution(pa)
fig_dist.update_layout(title="Monthly returns are right-skewed with contained losses")
fig_dist.update_xaxes(title_text="Monthly return")
for annotation, y_position in zip(fig_dist.layout.annotations, [0.92, 0.78], strict=False):
annotation.update(y=y_position, yshift=0)
fig_dist.show()
```
**Drawdown underwater.** Underwater curve over the strategy series;
reads against §4's max-drawdown number.
```python
fig_underwater = plot_drawdown_underwater(pa)
fig_underwater.update_layout(title="Underwater depth and time spent below the prior peak")
fig_underwater.show()
```
**Rolling Sharpe.** Twelve- and 36-month rolling Sharpe locate when the
realized Sharpe is paid out.
```python
fig_rolling = plot_rolling_sharpe(pa, windows=[12, 36])
for trace, label in zip(fig_rolling.data, ["12 months", "36 months"], strict=False):
trace.name = label
fig_rolling.update_layout(title="Performance persists across 12- and 36-month windows")
fig_rolling.show()
```
**Tail risk panel.** Value at risk and conditional value at risk at the two
conventional confidence levels, drawn against the empirical tail.
```python
fig_tail = plot_tail_risk_analysis(strat_arr)
_var_99 = float(np.quantile(strat_arr, 0.01))
for annotation, y_position in zip(
fig_tail.layout.annotations,
[0.96, 0.82, 0.68, 0.54],
strict=False,
):
annotation.update(y=y_position, yanchor="middle")
fig_tail.update_layout(
title="Monthly loss quantiles against the fitted normal tail",
height=650,
margin={"b": 70, "l": 60, "r": 30, "t": 80},
)
fig_tail.show()
```
## §5 Friction budget & cost sensitivity
Long-short equity at monthly cadence has limited turnover but doubled
cost exposure (long + short legs) and short-leg borrow costs. Two
layers cover the friction budget:
1. **Registry cost sweep** - the cost_sensitivity stage re-ran the selected
strategy at each level of the declared cost grid, holding everything else
fixed, so the curve below is the response of this one strategy to friction
rather than a spread across configurations. `setup.yaml` declares
`cost_sensitivity: 1`, so exactly one lineage is swept and the curve has one
Sharpe per level.
2. **Micro-cap-realistic extended grid** - because both long and short
legs concentrate in small-cap, illiquid names, asset-class-mean
spreads understate execution costs; the extended grid (up to 500
bps one-way) and turnover overlay surface the deployment-relevant
cost regime for the names actually selected.
```python
with sqlite3.connect(str(_db)) as _con:
cost_df = pl.DataFrame(
_con.execute(
"""
SELECT
b.spec_json,
bm.sharpe,
bm.sharpe_ci95_lo,
bm.sharpe_ci95_hi,
bm.max_drawdown
FROM backtest_runs b
JOIN backtest_metrics bm ON bm.backtest_hash = b.backtest_hash
JOIN prediction_sets p ON b.prediction_hash = p.prediction_hash
WHERE b.stage = 'cost_sensitivity'
AND p.split = 'validation'
AND b.prediction_hash = ?
AND bm.sharpe IS NOT NULL
AND (bm.num_trades IS NULL OR bm.num_trades > 0)
""",
(TOP_PHASH,),
).fetchall(),
schema=[
"spec_json",
"sharpe",
"sharpe_ci95_lo",
"sharpe_ci95_hi",
"max_drawdown",
],
orient="row",
)
```
```python
def _cost_bps(spec_str: str) -> float:
spec = json.loads(spec_str)
bc = spec.get("backtest_config", {})
comm = bc.get("commission", {}) or {}
slip = bc.get("slippage", {}) or {}
return float((comm.get("rate", 0) + slip.get("rate", 0)) * 10000)
cost_df = cost_df.with_columns(
pl.col("spec_json").map_elements(_cost_bps, return_dtype=pl.Float64).alias("cost_bps")
)
# Empty here means the cost stage swept a different lineage from the one selected
# above, which is a disagreement between two selection paths rather than a missing
# sweep, and it has to stop rather than draw an empty axis.
if cost_df.is_empty():
raise RuntimeError(
f"No cost_sensitivity run is registered against the selected prediction "
f"{TOP_PHASH}. The cost notebook swept a different lineage, so its curve does "
"not describe the strategy this section reports."
)
cost_curve = cost_df.select("cost_bps", "sharpe", "sharpe_ci95_lo", "sharpe_ci95_hi").sort(
"cost_bps"
)
print("Cost sensitivity of the selected strategy (validation):")
with pl.Config(tbl_rows=cost_curve.height):
print(cost_curve)
```
```python
cost_range = setup["costs"]["per_leg_cost_bps_range"]
fig, ax = plt.subplots(figsize=(9, 4))
xs = cost_curve["cost_bps"].to_numpy()
ax.fill_between(
xs,
cost_curve["sharpe_ci95_lo"].to_numpy(),
cost_curve["sharpe_ci95_hi"].to_numpy(),
alpha=0.18,
color=COLORS["slate"],
label="bootstrap confidence interval",
)
ax.plot(
xs,
cost_curve["sharpe"].to_numpy(),
color=COLORS["blue"],
linewidth=1.4,
label="Sharpe, net of the charge",
)
ax.axhline(0, color=COLORS["neutral"], linewidth=0.8, linestyle="--")
ax.axvspan(
cost_range[0],
cost_range[1],
color=COLORS["amber"],
alpha=0.10,
label=f"protocol per-leg cost ({cost_range[0]}-{cost_range[1]} bps)",
)
ax.set_xlabel("Per-leg cost (bps)")
ax.set_ylabel("Sharpe (validation)")
add_message_title(
ax,
"Sharpe declines slowly across the whole declared cost grid",
subtitle="Validation months; the strategy is unchanged, only what it pays to trade",
)
ax.legend(loc="best", fontsize=8, frameon=False)
show_with_alt(
fig,
"Validation Sharpe against the per-leg cost chargedПолный текст с указанием источника опубликован на условиях его лицензии. Лицензия: MIT
Это краткое изложение подготовлено исследовательским агентом Stratmill по оригиналу и не является его копией.