Comparer les stops de position FX sans biais de sélection
Résumé
Cette étude de cas compare des règles de stop fixes et suiveuses déclarées pour des positions FX. Elle sélectionne d’abord une stratégie parente par étiquette dans une cohorte de validation scellée, puis ne fait varier que la règle de risque de position. Comme les stops réagissent aux mouvements de prix pendant la période de détention, ils peuvent distinguer des trajectoires ayant le même rendement final. Le notebook rattache également les résultats de risque aux mêmes stratégies parentes que celles utilisées pour les comparaisons de coûts, afin que les différences entre étapes restent interprétables.
La méthodologie met l’accent sur les contrôles de recherche : exclure de la sélection des stratégies les lignes d’analyse de sensibilité aux coûts, vérifier que chaque comparaison ne modifie qu’une seule règle et figer l’ensemble complet des variantes avant de choisir un gagnant. Cela préserve les résultats perdants et indique le nombre d’essais ayant produit le meilleur résultat de validation. Les éléments présentés portent sur le cadre de comparaison en validation ; aucun chiffre de performance n’est fourni ici. Les effets des stops dépendent des barres observées et du modèle d’exécution, et choisir un stop sur la validation peut exagérer son bénéfice apparent. Un jeu de données de réserve devrait évaluer l’ensemble figé une seule fois. Cette étude de cas ne comporte ni drawdown au niveau du portefeuille ni limite de perte quotidienne ; ses contrôles s’appliquent donc aux positions individuellement plutôt qu’au portefeuille global.
Idées clés
- Une règle de stop réagit aux mouvements de prix pendant la période de détention, et pas seulement aux prix de début et de fin.
- Comparez les variantes de risque à partir d’une stratégie parente commune sélectionnée à l’avance afin que les résultats soient attribuables à la règle modifiée.
- Excluez les lignes d’analyse de sensibilité aux coûts de la sélection, car des hypothèses de coûts différentes peuvent fausser le classement des stratégies.
- Figez l’ensemble complet des variantes et conservez les variantes perdantes afin de rendre visibles le nombre d’essais et les résultats négatifs.
- Les améliorations observées en validation sont sujettes au biais de sélection et doivent être évaluées sur les données de réserve par rapport à l’ensemble figé.
Étiquettes
Texte intégral
# 15_risk_management.py
```py
# ---
# jupyter:
# jupytext:
# cell_metadata_filter: tags,-all
# text_representation:
# extension: .py
# format_name: percent
# format_version: '1.3'
# jupytext_version: 1.19.3
# kernelspec:
# display_name: Python 3 (ipykernel)
# language: python
# name: python3
# ---
# %% [markdown]
# # Position Risk Controls - FX Pairs
#
# The two stages before this one decided what to hold and how much. A position risk rule decides
# when to stop holding it. That makes this the first stage whose effect depends on the price path
# inside a holding period rather than on its endpoints: a position that ends the period down four
# percent and one that dipped nine percent on the way to ending down four are indistinguishable
# to every earlier stage, and a stop separates them.
#
# The rules compared here are declared in `config/setup.yaml`, not invented in the notebook: four
# fixed stop-loss levels and five trailing stops. Portfolio-level controls - a drawdown cap or a
# daily loss limit applied to the book as a whole - are absent for this case study, so nothing in
# this stage restrains the aggregate. Each position is governed on its own.
#
# This notebook selects one validation strategy per label from the equal-weight and allocation
# populations, then changes only its predeclared position-risk rule. Cost-sensitivity results are
# excluded from selection, and the reason is a selection defect rather than tidiness: a cost sweep
# re-prices one strategy under several assumptions, so a ranking that included those rows would
# advance whichever strategy had been measured under the kindest cost model. The comparison would
# read as a difference between strategies and be a difference between assumptions.
#
# **Learning objectives**
#
# - Select one parent strategy from an immutable validation cohort.
# - Compare declared position controls while preserving upstream identity.
# - Freeze the complete risk population before final strategy selection.
# - See why the count of variants tried belongs beside any improvement they produce.
#
# **Book reference**: Chapter 19
#
# **Prerequisite**: `14_portfolio_management`.
# %%
"""Run the declared FX position-risk controls for each label."""
from copy import deepcopy
from typing import Any
import polars as pl
import yaml
from case_studies.research import (
BacktestResult,
CandidateSet,
OfficialPopulation,
Result,
candidate_set_supersedes,
open_study,
plan_backtests,
population_supersedes,
research_name,
reuse_disclosure,
run_backtests,
superseded_members,
)
from case_studies.utils.backtest_presets import EngineBacktestConfig
from case_studies.utils.strategy_analysis import selectable_validation_candidates
from case_studies.utils.sweep_config import (
get_portfolio_risk_controls,
get_position_risk_controls,
get_top_n_predictions,
)
from utils.paths import get_case_study_dir
from utils.reproducibility import set_global_seeds
# %% tags=["parameters"]
CASE_STUDY_ID = "fx_pairs"
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""
LABEL = ""
SPLIT = "validation"
TOP_K = 0
TOP_N_PREDICTIONS = None
MAX_RISK_VARIANTS = 0
# How many parents per label the overlay grid sits on. `None` reads
# `backtest.sweep.top_n_predictions.risk_overlay`, which every case study declares as 1, and one
# is narrow on purpose: an overlay is a second search over the same validation folds, so the
# question the book asks is whether a control improves the configuration the funnel already
# chose. Until 2026-09-20 that 1 was a literal `min(eligible, ...)` below rather than a number
# read from the declaration, which left this case study unable to answer at any other width
# while four others could. A wider run changes the member list of every name published here, so
# it is narrowing and widening alike that need `POPULATION_NAME`, asserted below.
TOP_N_COMBOS = None
SEED = 42
RUN_SWEEP = True
FORCE_REBACKTEST = False
POPULATION_NAME = ""
# `POPULATION_NAME` scopes what this notebook *writes*. The equal-weight baselines it reads are
# written by `13_backtest` and carry the canonical name, so a scoped run that inherits the scope
# on that lookup asks for a population no run ever wrote and stops with "resolved to 0 current
# identities among 0 snapshots". Empty string reads the canonical baselines while the outputs
# stay scoped, which is the shape a partial re-run needs; None means "same scope as the outputs".
# `14_portfolio_management` gained the same parameter in #1109 and this is its sibling.
BASELINE_POPULATION_NAME: str | None = None
# `df20d72ab319` was the tip of `fx_pairs:risk-overlay-backtests` when it was written, and
# `create` accepts the tip and nothing else, so the literal was correct exactly until this
# notebook next published. `"live"` names the lineage and is resolved against it at run time.
SUPERSEDES_RISK_BACKTESTS: str = "live"
# A candidate set is immutable under its name, exactly as a population is, so a rebuilt upstream
# generation has to name the set it replaces. Keyed by the full set name because that is what the
# refusal prints: pasting back the name it names is the obvious thing to try, and it has to work.
# Resolved through `candidate_set_supersedes` rather than passed straight to `create`, because a
# reader's clean clone has no generation to supersede and `create` refuses a first version that
# claims to replace one.
#
# `"live"` names the lineage rather than a generation of it, which is what stops these going
# stale again. Three of the four hashes it replaces were already dead: the three
# `pre-risk-strategies` sets moved on 2026-09-09 (`d966caa61faf` -> `669d50f0f525`,
# `fde7af05fff6` -> `93b56a9dfb42`, `ff269dc95622` -> `f70630285818`), so each named the
# generation its own successor had replaced and the next membership move here would have been
# refused at the freeze, after the fit. `bf21ae4c9070` was still the head of
# `fx_pairs:holdout-candidates` and would have gone the same way on the next publish.
# See `case_studies.research.population.SUPERSEDES_LIVE`.
SUPERSEDES_CANDIDATE_SETS: dict[str, str] = {
"fx_pairs:fwd_ret_1d:pre-risk-strategies": "live",
"fx_pairs:fwd_ret_5d:pre-risk-strategies": "live",
"fx_pairs:fwd_ret_21d:pre-risk-strategies": "live",
"fx_pairs:holdout-candidates": "live",
}
# %% [markdown]
# ## Select the parent strategy
#
# Production uses the same signal-plus-allocation candidate cohort as the cost notebook. Preview
# mode resolves one deterministic allocation result from the reduced prediction catalog.
#
# One parent per label, chosen by best validation backtest Sharpe from the sealed pre-risk cohort.
# Sharing the cohort with the cost notebook is deliberate: both stages branch from the same
# strategy, so their results are siblings of one parent and can be read against each other. If
# each stage picked its own parent, a cost result and a risk result for the same label would
# describe two different strategies and the comparison between them would mean nothing.
#
# The cohort is sealed before it is ranked, which is what makes "best" a statement about a fixed
# set. A cohort that could still gain members after the selection would let a later run change
# which strategy this notebook chose, retroactively, with no record that it had changed.
#
# The `identity_status == "current"` filter below carries the same trap `13_backtest` documents:
# it names the schema version a row was written under, not whether its producer still publishes
# it. A refit leaves the replaced generation in the registry, complete and current, so the filter
# alone would let a retired prediction set into the cohort - and nothing would fail, because a
# retired prediction still produces a correct backtest. `superseded_members` reads the lineage,
# which is where the retirement is recorded.
# %% tags=["results"]
set_global_seeds(SEED)
universe_symbols = yaml.safe_load(
(get_case_study_dir(CASE_STUDY_ID) / "config" / "setup.yaml").read_text()
)["universe"]["symbols"]
n_assets = len(universe_symbols)
if SPLIT != "validation":
raise ValueError("risk comparison uses validation backtests")
if FORCE_REBACKTEST:
raise ValueError("identical complete backtests are reused by identity")
if not RUN_SWEEP:
raise ValueError("set RUN_SWEEP=True to execute the visible risk request")
study = open_study(CASE_STUDY_ID, execution_tier=EXECUTION_TIER, workspace=WORKSPACE or None)
# The execution tier decides which registry namespace this run reads and writes;
# the reduction knobs decide only how much of it is covered. Inferring the tier
# from the knobs conflated the two, so any reduced run went looking for preview
# predictions - and a reduced run over a canonical upstream, which is what the
# test suite exercises, then resolved no rows at all.
include_preview = EXECUTION_TIER == "preview"
def _resolve_baseline_scope(output_scope: str, input_scope: str | None) -> str:
return output_scope if input_scope is None else input_scope
baseline_population_name = _resolve_baseline_scope(POPULATION_NAME, BASELINE_POPULATION_NAME)
# The tier decides the namespace, so a canonical run may legitimately be narrowed -
# but a narrowed run declares a different set of members than the canonical
# population does, and a population is immutable once written. Such a run must
# publish under its own name rather than register a partial snapshot of the risk overlay sweep
# under the canonical one.
if (
(
TOP_K
or TOP_N_PREDICTIONS is not None
or MAX_RISK_VARIANTS
or LABEL
or TOP_N_COMBOS is not None
)
and not include_preview
and not POPULATION_NAME
):
raise ValueError(
"this run does not sweep the declared risk overlay population, so it cannot publish "
"the canonical one; pass POPULATION_NAME to give it its own"
)
if TOP_N_COMBOS is None:
TOP_N_COMBOS = get_top_n_predictions(CASE_STUDY_ID, "risk_overlay")
if TOP_N_COMBOS < 1:
raise ValueError("the risk overlay needs at least one parent per label")
catalog = study.predictions.table(include_preview=include_preview).filter(
(pl.col("identity_status") == "current")
& (pl.col("split") == SPLIT)
& pl.col("complete")
& (pl.col("execution_tier") == ("preview" if include_preview else "canonical"))
)
# `identity_status` is the schema version a row was written under, not a statement about which
# generation its producer publishes. A model notebook that refits leaves the generation it
# replaced in the registry, complete and current, so this filter alone would carry a retired
# prediction set into the sweep. `superseded_members` reads the lineage instead - see
# `13_backtest`, which drops the same set before it freezes the baseline population.
# `SUPERSEDES_RISK_BACKTESTS` is the sentinel `"live"`, so it names the lineage this run
# publishes under and `population_supersedes` resolves the generation in force at write time.
# It resolves to nothing until that name has a first generation, which is also what a reader's
# clean clone sees; after that, an upstream refit changes this population's member list and the
# registry refuses the write without the tip. `13_backtest` states the reasoning once.
retired = superseded_members(study, member_kind="prediction")
if retired:
catalog = catalog.filter(~pl.col("prediction_hash").is_in(list(retired)))
if LABEL:
catalog = catalog.filter(pl.col("label") == LABEL)
if TOP_N_PREDICTIONS is not None:
catalog = catalog.sort("label", "family", "config_name", "checkpoint_value").head(
TOP_N_PREDICTIONS
)
if catalog.is_empty():
raise RuntimeError("risk comparison resolved no complete prediction rows")
def _open_backtests(population: OfficialPopulation) -> list[BacktestResult]:
population.require_complete()
opened = [Result.open(study, value) for value in population.members]
if any(not isinstance(result, BacktestResult) for result in opened):
raise TypeError(f"population {population.name!r} contains a non-backtest result")
return [result for result in opened if isinstance(result, BacktestResult)]
def _label(result: BacktestResult) -> str:
return str(result.lineage()["training_spec"]["label"])
def _registered_preview_allocations() -> pl.DataFrame:
"""The allocation backtests an upstream preview registered.
One predicate, read once. Stating it twice - once to decide which labels are covered
and once to pick a label's leader - makes the two required to agree forever: tighten
one and the other admits a label whose backtests it then rejects, which is exactly the
"no preview allocation backtests are registered" failure this exists to prevent.
"""
return study.backtests.table(include_preview=True).filter(
(pl.col("stage") == "allocation")
& (pl.col("execution_tier") == "preview")
& (pl.col("identity_status") == "current")
& pl.col("complete")
)
def _preview_leader(rows: pl.DataFrame, registered_allocations: pl.DataFrame) -> BacktestResult:
"""One allocation backtest for this label, read from what 14 registered.
Rebuilding the identity restated two of the upstream run's choices - its `top_k` and
which allocator came first - from this notebook's own defaults. A preview reduces
each notebook independently, so those are guesses about another run's parameters,
and a guess that is wrong looks for a hash nothing wrote instead of reporting the
disagreement.
"""
registered = registered_allocations.filter(
pl.col("prediction_hash").is_in(rows.get_column("prediction_hash").implode())
)
if registered.is_empty():
raise RuntimeError(
"no preview allocation backtests are registered for this prediction "
"catalog; run 14_portfolio_management at the same reduction first"
)
row = registered.sort("family", "config_name", "checkpoint_value", "backtest_hash").row(
0, named=True
)
result = Result.open(study, row["backtest_hash"], include_preview=True)
if not isinstance(result, BacktestResult) or not result.complete:
raise RuntimeError("the deterministic preview allocation result is not complete")
return result
selected_by_label: dict[str, list[BacktestResult]] = {}
candidate_sets: dict[str, CandidateSet] = {}
if include_preview:
# The labels come from what the upstream preview registered, the same rule the
# canonical branch below follows. Enumerating this notebook's own catalog asks
# _preview_leader for labels 14_portfolio_management was configured not to allocate,
# and it raises "no preview allocation backtests are registered" for each - reporting
# a reduction the run was told to make as a missing upstream.
registered_allocations = _registered_preview_allocations()
covered = catalog.filter(
pl.col("prediction_hash").is_in(
registered_allocations.get_column("prediction_hash").implode()
)
)
if covered.is_empty():
raise RuntimeError(
"no preview allocation backtests cover this prediction catalog; "
"run 14_portfolio_management at the same reduction first"
)
for label in sorted(covered.get_column("label").unique()):
selected_by_label[label] = [
_preview_leader(covered.filter(pl.col("label") == label), registered_allocations)
]
else:
baselines = _open_backtests(
OfficialPopulation.one(
study,
name=research_name(
CASE_STUDY_ID, "equal-weight-baselines", scope=baseline_population_name
),
)
)
allocations = _open_backtests(
OfficialPopulation.one(
study,
name=research_name(CASE_STUDY_ID, "allocation-backtests", scope=POPULATION_NAME),
)
)
# The labels come from the upstream populations this run resolved, not from this
# notebook's own catalog. A narrowed upstream covers fewer labels than the catalog
# holds, and rebuilding the list locally reproduces the upstream narrowing by
# convention. Unscoped, the run publishes canonical names and the two must agree.
upstream = [*baselines, *allocations]
upstream_labels = sorted({_label(result) for result in upstream})
if not upstream_labels:
raise RuntimeError("the upstream populations carry no labels")
if not POPULATION_NAME and upstream_labels != sorted(catalog.get_column("label").unique()):
raise RuntimeError(
"the canonical upstream populations do not cover every label in the catalog: "
f"upstream {upstream_labels}, "
f"catalog {sorted(catalog.get_column('label').unique())}"
)
# Eligibility and order both come from `selectable_validation_candidates`, the function
# `resolve_solvent_carrier` ranks. Re-deriving them here swept the risk variants over a
# strategy the case study never publishes: the populations above are read whole, and nothing
# applies the retired-prediction test to them, so a backtest whose prediction a later refit
# superseded still won on raw Sharpe. `16_costs` carried the same defect.
_eligible_order = {
row["backtest_hash"]: position
for position, row in enumerate(
selectable_validation_candidates(CASE_STUDY_ID, labels=[LABEL] if LABEL else None)
)
}
for label in upstream_labels:
members = [result for result in upstream if _label(result) == label]
eligible = [result for result in members if result.hash in _eligible_order]
if not eligible:
raise RuntimeError(
f"none of the {len(members)} upstream backtests for {label} is selectable: "
"every one is retired on the backtest or the prediction side, or belongs to no "
"population its producer publishes. Re-run the validation stages rather than "
"sweeping risk controls over a strategy nothing reports."
)
_set_name = research_name(
CASE_STUDY_ID, f"{label}:pre-risk-strategies", scope=POPULATION_NAME
)
# The frozen set records the field the selection actually saw, so it holds the
# selectable members and not every row the three populations list.
candidates = CandidateSet.create(
study,
name=_set_name,
members=eligible,
supersedes=candidate_set_supersedes(
study, name=_set_name, declared=SUPERSEDES_CANDIDATE_SETS.get(_set_name)
),
)
candidate_sets[label] = candidates
leaders = sorted(eligible, key=lambda result: _eligible_order[result.hash])[:TOP_N_COMBOS]
if not all(isinstance(leader, BacktestResult) for leader in leaders):
raise TypeError("strategy selection did not return a backtest")
selected_by_label[label] = leaders
pl.DataFrame(
[
{
"label": label,
"backtest_hash": result.hash,
"prediction_hash": result.registry_record()["prediction_hash"],
"stage": result.registry_record()["stage"],
}
for label, results in selected_by_label.items()
for result in results
]
)
# %% [markdown]
# ## Plan and freeze exact risk siblings
#
# Every position rule is declared in `config/setup.yaml`. Portfolio-level controls are absent for
# this case study. The identity audit removes only the risk block and chapter label; the prediction,
# signal, allocation, configured costs, and execution contract must remain unchanged. Production
# freezes every expected risk identity before the first backtest is written.
#
# The audit is what makes each result a sibling rather than another strategy. Two backtests that
# differ in their stop and in nothing else can be subtracted; two that differ in their stop and
# their cost model cannot, and no field in the output would say which case you are looking at.
# Removing exactly the risk block and the chapter label, then requiring the remainder to match,
# is how that is established rather than assumed.
#
# How a stop is evaluated is worth knowing before reading its results. `StopLoss` and
# `TrailingStop` in `ml4t.backtest.risk` trigger on the bar's low or high where the price frame
# supplies a range, and fall back to the close where it does not, so the same threshold is a
# different rule depending on what the bars carry. The fill is a separate configured choice:
# filling at the stop price, at the bar's extreme, or at the close give materially different
# answers for one triggered stop, and the gap between them widens exactly in the volatile periods
# a stop is meant to handle. A stop's measured benefit is therefore partly a statement about the
# fill model, and that is a declared assumption rather than a property of the market.
# %% tags=["results"]
position_controls = get_position_risk_controls(CASE_STUDY_ID)
if get_portfolio_risk_controls(CASE_STUDY_ID):
raise RuntimeError("FX declares position controls only")
if MAX_RISK_VARIANTS:
position_controls = position_controls[:MAX_RISK_VARIANTS]
if not position_controls:
raise RuntimeError("the position-risk grid is empty")
def _risk_payload(control: dict[str, Any]) -> dict[str, Any]:
if control["type"] == "time_exit":
rule = {"type": control["type"], "bars": int(control["bars"])}
else:
rule = {"type": control["type"], "threshold": float(control["threshold"])}
return {"name": control["name"], "position_rules": [rule]}
def _catalog_row(result: BacktestResult) -> pl.DataFrame:
prediction_hash = result.registry_record()["prediction_hash"]
row = catalog.filter(pl.col("prediction_hash") == prediction_hash)
if row.height != 1:
raise RuntimeError(f"prediction {prediction_hash} resolved to {row.height} catalog rows")
return row
def _strategy_arguments(result: BacktestResult) -> dict[str, Any]:
strategy = result.spec()["strategy"]
return {
"signal": deepcopy(strategy["signal"]),
"allocation": deepcopy(strategy.get("allocation")),
"execution_mode": strategy.get("rebalance", {}).get("mode"),
}
def _non_risk_projection(spec: dict[str, Any]) -> dict[str, Any]:
projected = deepcopy(spec)
projected.pop("chapter", None)
projected.pop("_runtime_backtest_config", None)
projected.get("strategy", {}).pop("risk", None)
metadata = projected.get("backtest_config", {}).get("metadata")
if isinstance(metadata, dict):
metadata.pop("chapter", None)
# An absolute filesystem path, and already excluded from the identity hash by
# `_HASH_EXCLUDED_METADATA` for that reason. Comparing it here makes the notebook
# refuse its own siblings from any checkout but the one that registered the parents.
metadata.pop("preset_path", None)
# The parent allocation row was serialized by whatever engine version registered it and the
# risk result by the installed one, so a field `BacktestConfig` has since gained is absent on
# one side and present on the other while both describe the same strategy. `ml4t-backtest`
# 0.1.3 to 0.1.6 added `account.lock_notional_update_mode` and
# `position_sizing.share_rounding`, both previously implicit defaults the schema made
# explicit: `NEAREST` is the rounding the engine already did, and `POSITION_LEGS` only bites
# under a `lock_notional` short cash policy. Every fx_pairs parent predates them, so this
# comparison reported a moved strategy field on two names for one behaviour.
#
# Round-tripping both sides through the installed schema states the comparison in one
# vocabulary, so it answers what this notebook built rather than which engine wrote the row it
# is compared against, and it covers the next added field without naming it.
# `14_portfolio_management.py`, `16_costs.py` and `19_strategy_analysis.py` already do this;
# this projection is the one that was missed. `ensure_backtest_spec` deliberately does NOT
# round-trip, because there the result is hashed and a dropped unknown key would move an
# identity; here it is compared and discarded. Metadata is merged back over the serialized
# view because the dataclass pins a schema and drops keys it does not know.
config = projected.get("backtest_config", {})
if EngineBacktestConfig is not None and config:
original_metadata = dict(metadata) if isinstance(metadata, dict) else {}
rebuilt = EngineBacktestConfig.from_dict(config).to_dict()
rebuilt_metadata = dict(rebuilt.get("metadata") or {})
rebuilt_metadata.update(original_metadata)
rebuilt["metadata"] = rebuilt_metadata
projected["backtest_config"] = rebuilt
return projected
risk_jobs = []
for label, selected_results in selected_by_label.items():
for selected in selected_results:
arguments = _strategy_arguments(selected)
for control in position_controls:
risk = _risk_payload(control)
plan = plan_backtests(
study,
predictions=_catalog_row(selected),
signal=arguments["signal"],
allocation=arguments["allocation"],
risk=risk,
chapter="19",
execution_mode=arguments["execution_mode"],
)
if len(plan.members) != 1:
raise RuntimeError("a risk plan must contain exactly one backtest")
risk_jobs.append(
{
"label": label,
"selected": selected,
"arguments": arguments,
"risk": risk,
"risk_name": control["name"],
"backtest_hash": plan.expected_hashes[0],
}
)
planned_hashes = [job["backtest_hash"] for job in risk_jobs]
if len(planned_hashes) != len(set(planned_hashes)):
raise RuntimeError("two planned risk requests collapse to the same identity")
risk_population = None
if not include_preview:
risks_name = research_name(CASE_STUDY_ID, "risk-overlay-backtests", scope=POPULATION_NAME)
risk_population = OfficialPopulation.create(
study,
name=risks_name,
member_kind="backtest",
members=planned_hashes,
supersedes=population_supersedes(
study, name=risks_name, declared=SUPERSEDES_RISK_BACKTESTS
),
)
print(f"Frozen expected risk population: {risk_population.hash}")
# %% [markdown]
# ## Execute the frozen risk grid, and validate what was frozen
#
# The population is validated in the cell that fills it: the expected set was written down before
# the first member ran, and `require_complete` is what turns that declaration into a published
# result.
# %% tags=["results"]
# A sweep that recomputes everything and a sweep that recomputes nothing print the same summary
# unless the two are counted apart. `run_backtests` serves an identity that is already registered
# and complete instead of running it again, which is what makes a re-run affordable and what makes
# a bare member count say nothing about whether this run did any work.
#
# The runner already knows which it did and says so per member in `execution.diagnostics`, as
# `status` "reused" or "completed". Comparing against the registered hashes instead would be
# wrong in both directions: a registered-but-partial backtest is in that set, gets recomputed and
# would report as reused, and a preview re-run reads a table that excludes preview rows by default
# and would report every reused member as computed.
run_status: list[str] = []
risk_results: list[BacktestResult] = []
risk_rows = []
for job in risk_jobs:
selected = job["selected"]
arguments = job["arguments"]
execution = run_backtests(
study,
predictions=_catalog_row(selected),
signal=arguments["signal"],
allocation=arguments["allocation"],
risk=job["risk"],
chapter="19",
execution_mode=arguments["execution_mode"],
)
if len(execution.results) != 1:
raise RuntimeError("a risk request must produce exactly one backtest")
result = execution.results[0]
if result.hash != job["backtest_hash"]:
raise RuntimeError("a completed risk identity differs from the frozen plan")
if _non_risk_projection(result.spec()) != _non_risk_projection(selected.spec()):
raise RuntimeError("a risk sibling changed a non-risk strategy field")
if result.registry_record()["stage"] != "risk_overlay" or not result.complete:
raise RuntimeError("a risk result is incomplete or misclassified")
risk_results.append(result)
run_status.extend(entry["status"] for entry in execution.diagnostics)
risk_rows.append(
{
"label": job["label"],
"risk_name": job["risk_name"],
"backtest_hash": result.hash,
"prediction_hash": result.registry_record()["prediction_hash"],
}
)
served = run_status.count("reused")
print(
f"Risk overlays: {reuse_disclosure(len(risk_results) - served, served)}, "
f"{len(risk_results)} in the population"
)
if not include_preview:
if risk_population is None:
raise RuntimeError("the canonical risk population was not frozen before execution")
risk_population.require_complete()
print(f"Official risk population: {risk_population.hash}")
pl.DataFrame(risk_rows).sort("label", "risk_name")
# %% [markdown]
# ## Freeze the set the holdout will choose from
#
# Everything this case study has backtested on validation goes in: the equal-weight baselines, the
# allocation variants, and the risk overlays. The comparison contract names the fields every member
# must agree on, so a candidate fitted against different labels, features or folds cannot silently
# join a set the holdout will pick from. Only identities are printed here; what the selection is
# worth is `19_strategy_analysis`'s question.
#
# The size of this set is the part that no table above reports and that every number below depends
# on. Nine position rules were tried for each parent, on top of the allocators tried before them
# and the configurations tried before those. The best member wins by some margin over the rest, and
# part of that margin is simply the best of many draws on one validation period. Nothing in a
# result marks it: the Sharpe, the drawdown and the hit rate of the winning stop are all correctly
# computed for that stop, and a reader shown only the winner sees a strategy that looks better
# rather than a maximum taken over a set whose size is not on the page. This is why the count of
# members matters as much as the leader, why the set is frozen rather than pruned to the winner,
# and why the holdout is spent once against the whole set rather than against the leader alone.
#
# Freezing rather than pruning also keeps the negative results. A stop level that made things
# worse is a finding about this strategy, and it is only visible while the members that lost are
# still in the set beside the one that won.
# %%
if not include_preview:
_holdout_set_name = research_name(CASE_STUDY_ID, "holdout-candidates", scope=POPULATION_NAME)
holdout_candidates = CandidateSet.create(
study,
name=_holdout_set_name,
members=[*baselines, *allocations, *risk_results],
comparison_contract={"comparable_fields": ["label_artifact", "feature_artifacts", "cv"]},
supersedes=candidate_set_supersedes(
study,
name=_holdout_set_name,
declared=SUPERSEDES_CANDIDATE_SETS.get(_holdout_set_name),
),
)
selected = holdout_candidates.best_validation_sharpe()
print(f"Frozen holdout candidate set: {holdout_candidates.hash}")
print(f"Validation-selected backtest: {selected.hash}")
else:
print("Preview risk results remain outside official populations and candidate sets.")
# %% [markdown]
# ## Key takeaways
#
# - Risk variants descend from the same signal-plus-allocation selection cohort as cost variants,
# so a cost result and a risk result for one label describe the same parent strategy.
# - Each comparison changes one declared position-risk rule. The identity audit enforces that,
# because two results differing in a stop and in something else cannot be subtracted and nothing
# in the output would say so.
# - Cost-sensitivity rows are excluded from selection: including them would advance the strategy
# measured under the kindest cost assumption, not the best strategy.
# - A stop reads the path inside a holding period, which no earlier stage does. Its measured
# benefit depends on what the bars carry and on the configured fill model.
# - The final candidate set contains signal, allocation, and risk results across every label, and
# its size is the number of trials the eventual winner was drawn from.
#
# The improvement a stop shows here is measured on validation, and the stop level was chosen by
# looking at that same validation period. Whether it survives is the holdout's question, asked
# once, and it is allowed to answer no.
```Reproduit dans son intégralité avec attribution, conformément à la licence de la source. Licence: MIT
Ce résumé a été rédigé par l’agent de recherche de Stratmill à partir de la source originale ; il n’en est pas une copie.