वैलिडेशन और होल्डआउट साक्ष्य से FX पेयर्स रणनीति चुनना और उसका मूल्यांकन
सारांश
यह नोटबुक स्थिर वैलिडेशन उम्मीदवार सेट से चुने गए FX पेयर्स कॉन्फ़िगरेशन को फिर बनाती है, फिर उसके वैलिडेशन और होल्डआउट साक्ष्य का आकलन करती है। चयन में पात्र वैलिडेशन Sharpe का उच्चतम मान लिया जाता है और बराबरी पर बैकटेस्ट पहचान का नियतात्मक टाई-ब्रेकर लागू होता है; लागत के विकल्प उम्मीदवार फ़ील्ड में शामिल नहीं होते। होल्डआउट वंशावली का मिलान मॉडल परिवार, कॉन्फ़िगरेशन, लेबल और चेकपॉइंट से किया जाता है, क्योंकि वास्तविक रीट्रेन की प्रशिक्षण पहचान अलग होती है। विश्लेषण पहले से पंजीकृत होल्डआउट पूर्वानुमान और बैकटेस्ट पढ़ता है, उत्पत्ति जाँचता है और वैलिडेशन, होल्डआउट तथा उनके अंतर के लिए अंतराल अनुमान और युग्मित तुलना प्रस्तुत करता है।
लागत और जोखिम प्रभावों की जाँच नियंत्रित संबद्ध तुलनाओं से की जाती है, जबकि होल्डआउट परिणाम पहले के चयन को बदल नहीं सकते। नोटबुक पंजीकृत साक्ष्य को लगातार रिपोर्ट करने के लिए बनाई गई है और आवश्यक आर्टिफ़ैक्ट अनुपस्थित होने पर विफल होती है। इसके दावे चुने गए उम्मीदवार सेट और मापी गई अवधियों तक सीमित हैं: अनिश्चितता अंतराल में शून्य शामिल हो सकता है, और होल्डआउट चुने गए कॉन्फ़िगरेशन का मूल्यांकन है, रणनीतियों में से नया चयन करने का अवसर नहीं।
मुख्य विचार
- स्थिर वैलिडेशन उम्मीदवार सेट और Sharpe नियम चुने गए कॉन्फ़िगरेशन को तय करते हैं।
- होल्डआउट परिणाम वैलिडेशन प्रशिक्षण पहचान के बजाय कॉन्फ़िगरेशन विवरण से चुने गए मॉडल से जोड़े जाते हैं।
- रणनीति बदलावों का प्रभाव अलग करने के लिए लागत और जोखिम विकल्पों की नियंत्रित तुलना की जाती है।
- अंतराल और युग्मित साक्ष्य मापे गए प्रदर्शन को अनिश्चित अनुमानों से अलग करने में मदद करते हैं।
- होल्डआउट साक्ष्य पहले से चुनी गई रणनीति का आकलन करता है और उस चयन को बदलने के लिए उपयोग नहीं किया जा सकता।
टैग
पूरा पाठ
# 19_strategy_analysis.py
```py
# ---
# jupyter:
# jupytext:
# cell_metadata_filter: tags,-all
# text_representation:
# extension: .py
# format_name: percent
# format_version: '1.3'
# jupytext_version: 1.19.3
# kernelspec:
# display_name: Python 3 (ipykernel)
# language: python
# name: python3
# ---
# %% [markdown]
# # Strategy Analysis - FX Pairs
#
# This notebook reads back the selection made from the immutable validation candidate set and the
# holdout lineage that selection determines, and assesses both. Selection uses validation backtest
# Sharpe with the backtest identity as the deterministic tie-breaker. Cost sensitivity is excluded
# from selection, and nothing measured on the holdout can revise the choice - not because a lock
# forbids it, but because the choice was made upstream against a set that is already frozen.
#
# The holdout results are produced by `17_holdout_predictions` and `18_holdout_backtest`. This
# notebook writes nothing; it fails if either is missing rather than producing them itself, so
# reading the holdout and deciding what to run against it stay separate acts.
#
# **Learning objectives**
#
# - Reproduce validation selection from an immutable, complete candidate set.
# - Verify that the holdout lineage on record is the one the selection determines.
# - Interpret cost and risk variants through controlled sibling comparisons.
# - Assess validation and holdout performance with interval and paired evidence.
#
# **Book reference**: Chapters 16-20
#
# **Prerequisites**: `17_holdout_predictions` and `18_holdout_backtest`, and the candidate set
# `15_risk_management` freezes.
# %%
"""Read back the selected FX validation lineage and its holdout, and assess both."""
import json
import sqlite3
import warnings
from collections import Counter
from copy import deepcopy
from typing import Any
import matplotlib.pyplot as plt
import numpy as np
import plotly.express as px
import polars as pl
import yaml
from case_studies.research import (
BacktestResult,
CandidateSet,
OfficialPopulation,
PredictionResult,
Result,
TrainingResult,
open_study,
)
from case_studies.research.holdout import build_holdout_training_spec
from case_studies.research.strategy import strategy_warmup_periods
from case_studies.utils.artifact_digest import value_digest
from case_studies.utils.backtest_loaders import get_backtest_config, load_backtest_prices_for
from case_studies.utils.backtest_presets import (
EngineBacktestConfig,
cost_view,
ensure_backtest_spec,
)
from case_studies.utils.backtest_runner import resolved_allow_short_selling
from case_studies.utils.cohort_metrics import compute_and_register
from case_studies.utils.paired_metrics import populate_paired_metrics
from case_studies.utils.registry import (
backtest_run_status,
load_backtest_metrics,
load_paired_metrics,
)
from case_studies.utils.registry.specs import training_hash_from_spec
from case_studies.utils.strategy_analysis import (
resolve_canonical_rank1_lineage,
resolve_solvent_carrier,
selectable_validation_candidates,
)
from case_studies.utils.uncertainty import ENTIRE_REGISTRY
from utils.paths import get_case_study_dir
from utils.style import COLORS
# %% tags=["parameters"]
CASE_STUDY_ID = "fx_pairs"
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""
CANDIDATE_SET_NAME = "fx_pairs:holdout-candidates"
# %% [markdown]
# ## Resolve the selected configuration and its holdout lineage
#
# Selection is not a parameter and is not made here. `resolve_solvent_carrier` reads the
# highest-Sharpe registered validation backtest across the baseline, allocation and risk-overlay
# stages, among runs that stayed solvent and belong to a generation still in force. Cost
# siblings are not candidates: a cost variant is a descendant of a selection rather than an
# entrant in one. Nothing on this page can revise the choice.
#
# The holdout lineage is matched to that selection by CONFIGURATION - family, configuration name,
# label and checkpoint - rather than by the validation model's training hash. A genuine retrain
# does not share that hash; that is what makes it a retrain, and a lineage query keyed on it can
# only ever find a validation fit scored over a later window.
#
# There is no lock and no ledger. The whole rule is: take the configuration validation ranked
# first, retrain it on everything up to the holdout window, predict, and run that same backtest
# configuration on the result. An earlier design pre-registered the lineage in a research lock,
# which made the holdout a one-shot transaction and therefore impossible to correct - any fix
# upstream needed a retrain the lock forbade, and the lock could not be reissued.
# %% tags=["results"]
study = open_study(CASE_STUDY_ID, execution_tier=EXECUTION_TIER, workspace=WORKSPACE or None)
# The selection is the highest-Sharpe registered validation backtest across the baseline,
# allocation and risk-overlay stages, among runs that stayed solvent and belong to a generation
# still in force. Read from the registry, so this page cannot report a configuration the
# validation stages did not rank first.
# The frozen set is the field, exactly as in `17` and `18`. This page reports that configuration
# as the case study's answer, so it resolves over the same restricted field rather than over the
# whole registry - otherwise a row the set never admitted could still move the common-support
# intersection and change which admitted row is reported.
holdout_candidates = CandidateSet.one(study, name=CANDIDATE_SET_NAME)
ADMITTED = frozenset(holdout_candidates.members)
# The prediction sets behind the admitted backtests. `compute_and_register` scopes cohorts by
# prediction rather than by backtest, and left unscoped it computes the effective trial count and
# the deflated Sharpe over the whole registry - every retired generation and every configuration
# no population publishes included. K is what the deflation divides by, so an unscoped call
# reports a correction computed over a population this page does not describe.
with sqlite3.connect(str(get_case_study_dir(CASE_STUDY_ID) / "run_log" / "registry.db")) as _con:
ADMITTED_PREDICTIONS = [
row[0]
for row in _con.execute(
"SELECT DISTINCT prediction_hash FROM backtest_runs "
f"WHERE backtest_hash IN ({','.join('?' * len(ADMITTED))})",
tuple(sorted(ADMITTED)),
)
]
carrier = resolve_solvent_carrier(CASE_STUDY_ID, admitted=ADMITTED)
selected_validation = study.results.open(str(carrier["val_backtest_hash"]))
selected_prediction = study.results.open(str(carrier["val_prediction_hash"]))
selected_training = study.results.open(str(carrier["training_hash"]))
selected_record = selected_validation.registry_record()
selected_prediction_record = selected_prediction.registry_record()
selected_training_record = selected_training.registry_record()
selected_training_spec = selected_training.spec()
selected_computation = selected_training_spec.get("computation", selected_training_spec)
# Everything this case study backtested on validation and still publishes: the equal-weight
# baselines, the allocation variants and the risk overlays. Cost siblings are excluded, because
# a cost variant is a descendant of a selection rather than a candidate for one.
#
# `selectable_validation_candidates` is the function the selection ranks, called here
# with the same frozen set, so the table below and the selection are the same field by
# construction rather than by two filters that have to agree. Rebuilding the field in a query
# beside it did not agree: that query excluded the *retired* set, which is not the *published*
# set. A backtest no population ever listed was retired by nobody, so an exclusion filter admits
# it while the membership test the selection applies does not - and one such row, `56070f34dff1`,
# sorted above the selected configuration on raw Sharpe and made this page refuse to render.
#
# The order is the resolver's own. Where a conformal candidate is in the field it re-ranks every
# member on the timestamps they all price, because a calibration that abstains through its
# warm-up books those decisions as zero and is otherwise compared against allocators measured
# over a longer span. Each row therefore carries both numbers: `sharpe` as the registry stored
# it, and `comparison_sharpe` as the selection read it.
candidate_field = selectable_validation_candidates(CASE_STUDY_ID, admitted=ADMITTED)
pl.DataFrame(
{
"field": ["selected backtest", "selected stage", "candidates ranked", "validation Sharpe"],
"value": [
selected_validation.hash,
str(carrier["val_stage"]),
str(len(candidate_field)),
f"{carrier['val_sharpe']:.4f}",
],
}
)
# %% [markdown]
# ## The selected validation lineage
#
# Model family, configuration, label, checkpoint and the data artifacts the fit read are printed
# together, because they are what the holdout comparison holds fixed. The provenance fields are
# checked rather than displayed: a selected training run with no recorded source commit or runtime
# cannot be reproduced by a reader, and a holdout number from a run nobody can reproduce is not
# evidence of anything.
# %% tags=["results"]
for field in ("label_artifact", "feature_artifacts", "cv"):
if not selected_computation.get(field):
raise ValueError(f"the selected training run records no {field}")
if not selected_training_record.get("git_commit"):
raise ValueError("the selected training run records no source commit")
if not json.loads(selected_training_record.get("runtime_json") or "{}"):
raise ValueError("the selected training run records no runtime provenance")
selected_identity = pl.DataFrame(
{
"field": [
"label",
"family",
"configuration",
"checkpoint kind",
"checkpoint value",
"training hash",
"prediction hash",
"validation backtest hash",
"source commit",
],
"value": [
str(selected_training_spec["label"]),
str(selected_training_spec["family"]),
str(selected_training_spec["config_name"]),
str(selected_prediction_record["checkpoint_kind"]),
str(selected_prediction_record["checkpoint_value"]),
selected_training.hash,
selected_prediction.hash,
selected_validation.hash,
str(selected_training_record["git_commit"]),
],
}
)
selected_identity
# %% [markdown]
# ## Validation candidate evidence
#
# Every candidate remains visible below. The displayed order reproduces the selection rule; IC and
# every metric other than Sharpe remain descriptive.
# %% tags=["results"]
def _metric_row(result: BacktestResult) -> dict[str, Any]:
metrics = load_backtest_metrics(
CASE_STUDY_ID,
backtest_hash=result.hash,
case_dir=study.root,
)
if metrics.height != 1:
raise ValueError(f"backtest {result.hash} has {metrics.height} metric rows")
return metrics.row(0, named=True)
candidate_rows = []
for _candidate in candidate_field:
member_hash = _candidate["backtest_hash"]
result = Result.open(study, member_hash)
if not isinstance(result, BacktestResult) or not result.complete:
raise ValueError(f"candidate {member_hash} is not a complete backtest")
record = result.registry_record()
lineage = result.lineage()
training = lineage["training_spec"]
prediction = Result.open(study, record["prediction_hash"])
if not isinstance(prediction, PredictionResult):
raise TypeError(f"candidate {member_hash} does not reference a prediction")
prediction_record = prediction.registry_record()
metric = _metric_row(result)
if record["stage"] not in {"signal", "allocation", "risk_overlay"}:
raise ValueError(f"candidate {member_hash} has an ineligible stage")
if any(metric.get(name) is None for name in ("sharpe", "sharpe_ci95_lo", "sharpe_ci95_hi")):
raise ValueError(f"candidate {member_hash} lacks Sharpe interval evidence")
candidate_rows.append(
{
"backtest_hash": result.hash,
"prediction_hash": record["prediction_hash"],
"stage": record["stage"],
"label": training["label"],
"family": training["family"],
"config_name": training["config_name"],
"checkpoint_kind": prediction_record["checkpoint_kind"],
"checkpoint_value": prediction_record["checkpoint_value"],
"sharpe": metric["sharpe"],
"comparison_sharpe": _candidate["comparison_sharpe"],
"sharpe_ci95_lo": metric["sharpe_ci95_lo"],
"sharpe_ci95_hi": metric["sharpe_ci95_hi"],
}
)
# Not re-sorted: the rows arrive in the order the selection ranked them, so the top row is the
# selected configuration by construction. The check below is kept anyway and is now a real one -
# it fails if the resolver ever returns a field whose first member is not the configuration it
# resolved, which would mean the two halves of the same function had come apart.
candidate_evidence = pl.DataFrame(candidate_rows)
if candidate_evidence["backtest_hash"][0] != selected_validation.hash:
raise ValueError(
f"the ranked candidate field begins with {candidate_evidence['backtest_hash'][0]} "
f"and the resolved carrier is {selected_validation.hash}. Both come from "
"`selectable_validation_candidates` over the same admitted set, so the ranking and "
"the resolution disagree with each other."
)
candidate_evidence
# %% tags=["results"]
candidate_figure = px.scatter(
candidate_evidence,
x="sharpe",
y="family",
color="stage",
facet_row="label",
hover_data=["config_name", "backtest_hash"],
title="Validation Sharpe distribution for the immutable FX candidate set",
labels={"sharpe": "Validation Sharpe", "family": "Model family"},
)
candidate_figure.add_vline(x=0, line_dash="dot", line_color=COLORS["recede"])
candidate_figure.show()
# %% [markdown]
# ## Controlled cost and risk comparisons
#
# Cost sensitivity and risk overlays are read from their official populations. A sibling enters a
# comparison only when its prediction, signal, allocation, rebalance, execution, and price identity
# match the selected lineage after removing the field being varied.
# %% tags=["results"]
def _comparison_projection(
result: BacktestResult,
*,
omit_costs: bool,
omit_risk: bool,
) -> dict[str, Any]:
projected = deepcopy(result.spec())
projected.pop("chapter", None)
projected.pop("_runtime_backtest_config", None)
if omit_risk:
projected.get("strategy", {}).pop("risk", None)
metadata = projected.get("backtest_config", {}).get("metadata")
if isinstance(metadata, dict):
metadata.pop("chapter", None)
# Absolute filesystem path, excluded from the identity hash by
# `case_studies/utils/registry/specs.py` for the same reason it is excluded here:
# comparing it makes the projection depend on which checkout wrote the row.
metadata.pop("preset_path", None)
if omit_costs:
config = projected.get("backtest_config", {})
config.pop("commission", None)
config.pop("slippage", None)
# The rows being compared were serialized by whatever engine version registered each one, so a
# field `BacktestConfig` has since gained is present on one side and absent on the other while
# both describe the same strategy. `ml4t-backtest` 0.1.3 to 0.1.6 added
# `account.lock_notional_update_mode` and `position_sizing.share_rounding`; every row in this
# registry written before 2026-09-12 lacks both, so a cost sibling written after it matched no
# selected strategy and this notebook stopped at "no controlled cost siblings match". Round
# -tripping both sides through the installed schema states the comparison in one vocabulary, so
# it answers whether two runs describe the same strategy rather than which engine wrote them,
# and it covers the next added field without naming it. `ensure_backtest_spec` deliberately does
# not round-trip, because its result is hashed and a dropped unknown key would move an identity;
# here the result is compared and discarded. Metadata is merged back over the serialized view
# because the dataclass pins a schema and drops keys it does not know.
config = projected.get("backtest_config", {})
if EngineBacktestConfig is not None and config:
original_metadata = dict(metadata) if isinstance(metadata, dict) else {}
rebuilt = EngineBacktestConfig.from_dict(config).to_dict()
if omit_costs:
rebuilt.pop("commission", None)
rebuilt.pop("slippage", None)
rebuilt_metadata = dict(rebuilt.get("metadata") or {})
rebuilt_metadata.update(original_metadata)
rebuilt["metadata"] = rebuilt_metadata
projected["backtest_config"] = rebuilt
return {
"prediction_hash": result.registry_record()["prediction_hash"],
"spec": projected,
}
cost_population = OfficialPopulation.one(study, name=f"{CASE_STUDY_ID}:cost-sensitivity-backtests")
risk_population = OfficialPopulation.one(study, name=f"{CASE_STUDY_ID}:risk-overlay-backtests")
cost_population.require_complete()
risk_population.require_complete()
selected_cost_core = _comparison_projection(selected_validation, omit_costs=True, omit_risk=True)
selected_risk_core = _comparison_projection(selected_validation, omit_costs=False, omit_risk=True)
cost_rows = []
for member_hash in cost_population.members:
result = Result.open(study, member_hash)
if not isinstance(result, BacktestResult):
raise TypeError("cost population contains a non-backtest result")
if _comparison_projection(result, omit_costs=True, omit_risk=True) != selected_cost_core:
continue
costs = cost_view(result.spec())
metric = _metric_row(result)
if any(metric.get(name) is None for name in ("sharpe", "sharpe_ci95_lo", "sharpe_ci95_hi")):
raise ValueError(f"cost sibling {result.hash} lacks Sharpe interval evidence")
cost_rows.append(
{
"total_cost_bps": costs["commission_bps"] + costs["slippage_bps"],
"sharpe": metric["sharpe"],
"sharpe_ci95_lo": metric["sharpe_ci95_lo"],
"sharpe_ci95_hi": metric["sharpe_ci95_hi"],
"backtest_hash": result.hash,
}
)
if not cost_rows:
raise ValueError("no controlled cost siblings match the selected strategy")
cost_evidence = pl.DataFrame(cost_rows).sort("total_cost_bps")
cost_evidence
# %% tags=["results"]
cost_figure = px.line(
cost_evidence,
x="total_cost_bps",
y="sharpe",
markers=True,
title="Validation Sharpe for exact cost siblings of the selected FX strategy",
labels={"total_cost_bps": "Total cost per traded leg (basis points)", "sharpe": "Sharpe"},
)
cost_figure.add_hline(y=0, line_dash="dot", line_color=COLORS["recede"])
cost_figure.show()
# %% tags=["results"]
risk_rows = []
for member_hash in risk_population.members:
result = Result.open(study, member_hash)
if not isinstance(result, BacktestResult):
raise TypeError("risk population contains a non-backtest result")
if _comparison_projection(result, omit_costs=False, omit_risk=True) != selected_risk_core:
continue
metric = _metric_row(result)
if any(metric.get(name) is None for name in ("sharpe", "sharpe_ci95_lo", "sharpe_ci95_hi")):
raise ValueError(f"risk sibling {result.hash} lacks Sharpe interval evidence")
risk = result.spec()["strategy"]["risk"]
risk_rows.append(
{
"risk_name": risk["name"],
"sharpe": metric["sharpe"],
"sharpe_ci95_lo": metric["sharpe_ci95_lo"],
"sharpe_ci95_hi": metric["sharpe_ci95_hi"],
"backtest_hash": result.hash,
}
)
if not risk_rows:
raise ValueError("no controlled risk siblings match the selected strategy")
risk_evidence = pl.DataFrame(risk_rows).sort("sharpe", descending=True)
risk_evidence
# %% tags=["results"]
risk_figure = px.bar(
risk_evidence.sort("sharpe"),
x="sharpe",
y="risk_name",
orientation="h",
title="Validation Sharpe for exact risk siblings of the selected FX strategy",
labels={"sharpe": "Sharpe", "risk_name": "Position-risk rule"},
)
risk_figure.add_vline(x=0, line_dash="dot", line_color=COLORS["recede"])
risk_figure.show()
# %% [markdown]
# The cost curve is allowed to be monotone, nonmonotone, entirely positive, or entirely negative.
# The computed summary reports the observed shape without requiring a crossing. Risk rows are ordered
# descriptively and do not replace the candidate-set selection rule.
# %% tags=["results"]
cost_differences = cost_evidence.get_column("sharpe").diff().drop_nulls()
if (cost_differences < 0).all():
cost_shape = "strictly decreasing over the declared grid"
elif (cost_differences <= 0).all():
cost_shape = "nonincreasing over the declared grid"
else:
cost_shape = "nonmonotone over the declared grid"
if (cost_evidence.get_column("sharpe") > 0).all():
sign_summary = "all observed Sharpe estimates are positive"
elif (cost_evidence.get_column("sharpe") < 0).all():
sign_summary = "all observed Sharpe estimates are negative"
else:
sign_summary = "the observed Sharpe estimates include both signs or zero"
controlled_summary = pl.DataFrame(
{
"comparison": ["cost sensitivity", "position-risk controls"],
"finding": [
f"Sharpe is {cost_shape}; {sign_summary}.",
(
f"The highest point estimate belongs to {risk_evidence['risk_name'][0]}; "
"the interval columns determine whether that ordering is resolved."
),
],
}
)
controlled_summary
# %% [markdown]
# ## Require the holdout lineage the selected configuration determines
#
# The holdout results are not whatever happens to carry the holdout split; they are required to
# be the ones the selected configuration determines. The match is not on its validation training
# hash - a genuine retrain never shares the validation model's training identity, which is the
# whole point of a retrain - and it is not on family, configuration name and label either, because
# several training specifications carry the same three names. The holdout training identity is
# derived here the way `17_holdout_predictions` derives it, and the query asks for that hash.
# %% tags=["results"]
_carrier_ck = selected_prediction_record["checkpoint_kind"]
_carrier_cv = selected_prediction_record["checkpoint_value"]
_observation_timeline = (
pl.read_parquet(study.root / "labels" / f"{carrier['label']}.parquet")
.get_column("timestamp")
.unique()
.sort()
.to_list()
)
EXPECTED_HOLDOUT_TRAINING_HASH = training_hash_from_spec(
build_holdout_training_spec(
study,
study.results.open(carrier["training_hash"]).spec(),
timeline=_observation_timeline,
case_study=CASE_STUDY_ID,
)
)
with sqlite3.connect(str(study.root / "run_log" / "registry.db")) as _conn:
_holdout_rows = _conn.execute(
"""
SELECT prediction_hash
FROM prediction_sets
WHERE split = 'holdout'
AND training_hash = ?
AND checkpoint_kind IS ? AND checkpoint_value IS ?
ORDER BY prediction_hash
""",
(EXPECTED_HOLDOUT_TRAINING_HASH, _carrier_ck, _carrier_cv),
).fetchall()
if len(_holdout_rows) != 1:
raise ValueError(
f"holdout training run {EXPECTED_HOLDOUT_TRAINING_HASH} resolves "
f"{len(_holdout_rows)} prediction sets at checkpoint {_carrier_ck}={_carrier_cv}; "
"exactly one is required - run 17_holdout_predictions, or delete the superseded one"
)
holdout_prediction = Result.open(study, _holdout_rows[0][0])
_holdout_backtests = (
study.backtests.table()
.filter((pl.col("prediction_hash") == holdout_prediction.hash) & pl.col("complete"))
.get_column("backtest_hash")
.to_list()
)
if len(_holdout_backtests) != 1:
raise ValueError(
f"holdout prediction {holdout_prediction.hash} has {len(_holdout_backtests)} complete "
"backtests; exactly one is required - run 18_holdout_backtest"
)
holdout_backtest = Result.open(study, _holdout_backtests[0])
# One complete backtest on the holdout prediction is not the same as the selection's replay
# having been the one that produced it, and comparing the `strategy` block alone does not
# close the gap: commissions, slippage, account settings and the price identity all sit
# outside that block, so a cost variant or a stale-price run registered against the same
# prediction set would pass a strategy comparison and be reported as the selection's holdout.
#
# The backtest hash covers every input that changes the result, by construction. So the
# expected specification is rebuilt here exactly as `18_holdout_backtest` builds it - the
# selection's registered spec, the holdout prediction, the holdout price frame with the
# strategy's declared warmup, and the digest of that frame - and the registered backtest is
# required to BE that identity rather than to resemble it.
_holdout_prices = load_backtest_prices_for(
CASE_STUDY_ID,
str(carrier["label"]),
split="holdout",
warmup_periods=strategy_warmup_periods(json.loads(selected_record["spec_json"])),
)
_expected_spec = ensure_backtest_spec(
CASE_STUDY_ID,
get_backtest_config(CASE_STUDY_ID),
json.loads(selected_record["spec_json"]),
prices=_holdout_prices,
prediction_hash=holdout_prediction.hash,
initial_cash=get_backtest_config(CASE_STUDY_ID).initial_cash,
)
_expected_spec["chapter"] = "ch20"
_expected_spec.setdefault("input_identity", {})["prices"] = value_digest(_holdout_prices)
_expected_spec["backtest_config"]["account"]["allow_short_selling"] = resolved_allow_short_selling(
_expected_spec, None
)
EXPECTED_HOLDOUT_BACKTEST_HASH = backtest_run_status(
CASE_STUDY_ID, holdout_prediction.hash, _expected_spec
).backtest_hash
if holdout_backtest.hash != EXPECTED_HOLDOUT_BACKTEST_HASH:
raise ValueError(
f"the registered holdout backtest is {holdout_backtest.hash}, but replaying the "
f"carrier on the holdout produces {EXPECTED_HOLDOUT_BACKTEST_HASH}. The registered "
"run is a different configuration, not this carrier's holdout - re-run "
"18_holdout_backtest rather than reporting it."
)
holdout_training = Result.open(study, holdout_prediction.registry_record()["training_hash"])
if not isinstance(holdout_training, TrainingResult) or not holdout_training.complete:
raise ValueError("the holdout training result is incomplete")
if not isinstance(holdout_prediction, PredictionResult) or not holdout_prediction.complete:
raise ValueError("the holdout prediction result is incomplete")
if not isinstance(holdout_backtest, BacktestResult) or not holdout_backtest.complete:
raise ValueError("the holdout backtest result is incomplete")
if any(
result.execution_tier != "canonical"
for result in (holdout_training, holdout_prediction, holdout_backtest)
):
raise ValueError("holdout lineage must use canonical execution")
# A holdout training identity equal to the validation one means the re-keying changed nothing:
# the model is a validation fit predicting forward rather than one trained up to the window.
if holdout_training.hash == selected_training.hash:
raise ValueError(
f"the holdout training identity equals the validation one ({selected_training.hash}), "
"so the holdout model was not refitted"
)
if not holdout_backtest.spec().get("input_identity", {}).get("prices"):
raise ValueError("the holdout backtest lacks canonical price identity")
_holdout_fold = holdout_training.spec()["computation"]["cv"]["folds"][-1]
holdout_identity = pl.DataFrame(
{
"field": [
"holdout training",
"holdout prediction",
"holdout backtest",
"holdout train window",
"holdout evaluation window",
],
"value": [
holdout_training.hash,
holdout_prediction.hash,
holdout_backtest.hash,
f"{_holdout_fold['train_start']} to {_holdout_fold['train_end']}",
f"{_holdout_fold['val_start']} to {_holdout_fold['val_end']}",
],
}
)
holdout_identity
# %% [markdown]
# ## Validation and holdout evidence
#
# Point estimates and intervals are displayed by exact selected identity. Statistical comparisons use
# registered paired evidence. The holdout may disconfirm the validation result and cannot trigger
# fallback or reselection.
# %% tags=["results"]
required_metrics = {
"sharpe",
"sharpe_ci95_lo",
"sharpe_ci95_hi",
"total_return",
"max_drawdown",
"max_dd_ci95_lo",
"max_dd_ci95_hi",
"volatility",
"avg_turnover",
"num_trades",
}
performance_rows = []
for period, result in (("validation", selected_validation), ("holdout", holdout_backtest)):
metrics = _metric_row(result)
if any(
metrics.get(name) is None or not np.isfinite(metrics[name]) for name in required_metrics
):
raise ValueError(f"{period} result lacks finite performance evidence")
performance_rows.append(
{
"period": period,
"backtest_hash": result.hash,
**{name: metrics[name] for name in required_metrics},
}
)
selected_performance = pl.DataFrame(performance_rows)
selected_performance
# %% tags=["results"]
performance_figure = px.bar(
selected_performance,
x="period",
y="sharpe",
error_y=selected_performance.get_column("sharpe_ci95_hi")
- selected_performance.get_column("sharpe"),
error_y_minus=selected_performance.get_column("sharpe")
- selected_performance.get_column("sharpe_ci95_lo"),
title="Selected FX strategy Sharpe in validation and holdout windows",
labels={"period": "Window", "sharpe": "Sharpe"},
)
performance_figure.add_hline(y=0, line_dash="dot", line_color=COLORS["recede"])
performance_figure.show()
# %% [markdown]
# ### Register the cohort and paired evidence this section reads
#
# The bootstrapped comparisons and the effective-rank cohort statistics are computed here rather
# than assumed. They used to be a side effect of the holdout lock transaction; with that gone,
# the notebook that reads them is the notebook that has to produce them.
#
# The selected configuration is passed in rather than left to the populator. Left to itself it
# ranks the registry on raw Sharpe, which would be a second selector sitting beside
# `resolve_solvent_carrier` and the cost sweep - and a raw ranking has no notion of a retired
# generation, so it would pair the superseded conformal-v2 backtest and describe a configuration
# this case study does not report.
# %% tags=["results"]
_periods_per_year = int(
yaml.safe_load((get_case_study_dir(CASE_STUDY_ID) / "config" / "setup.yaml").read_text())[
"evaluation"
]["periods_per_year"]
)
# The cohort statistics warn per metric for every cohort whose leader holds a constant position
# - a deflated Sharpe needs return variance and a correlation matrix, and neither is defined
# there. The warnings are collected rather than printed: each one carries the absolute path of
# the module that raised it, which is this machine's path and does not belong in a committed
# notebook. What they mean is a count, so a count is what is reported.
with warnings.catch_warnings(record=True) as _cohort_warnings:
warnings.simplefilter("always")
_cohort_counts = compute_and_register(CASE_STUDY_ID, prediction_hashes=ADMITTED_PREDICTIONS)
_undefined = Counter(str(entry.message).split(" for ")[0] for entry in _cohort_warnings)
# The cohort call above is scoped to `ADMITTED_PREDICTIONS` and this one is not: the pairs
# are selected from every registered prediction set. Stated rather than defaulted;
# narrowing it would change published numbers, so it is a separate decision from this line.
_paired_rows = populate_paired_metrics(
CASE_STUDY_ID,
periods_per_year=_periods_per_year,
carrier=resolve_canonical_rank1_lineage(CASE_STUDY_ID, admitted=ADMITTED),
replace_all=True,
prediction_hashes=ENTIRE_REGISTRY,
)
print(
f"cohort_metrics: {sum(_cohort_counts[k] for k in ('family', 'stagelabel', 'label'))} rows; "
f"backtest_paired_metrics: {sum(1 for row in _paired_rows if 'skip' not in row)} pairs"
)
for _metric, _count in sorted(_undefined.items()):
print(f" undefined on a constant-position cohort: {_metric} x{_count}")
# %% tags=["results"]
holdout_pairs = load_paired_metrics(
CASE_STUDY_ID,
challenger_hash=holdout_backtest.hash,
case_dir=study.root,
)
validation_pairs = load_paired_metrics(
CASE_STUDY_ID,
challenger_hash=selected_validation.hash,
case_dir=study.root,
)
if holdout_pairs.is_empty() or validation_pairs.is_empty():
raise ValueError("required paired validation or holdout evidence is missing")
validation_to_holdout = holdout_pairs.filter(
(pl.col("benchmark_hash") == selected_validation.hash)
& (pl.col("benchmark_kind") == "val_rank1_self")
)
holdout_to_benchmark = holdout_pairs.filter(
pl.col("benchmark_kind") == "equal_weight_holdout_side_artifact"
)
validation_to_benchmark = validation_pairs.filter(
pl.col("benchmark_kind") == "equal_weight_side_artifact"
)
paired_required = {
"sharpe_diff",
"sharpe_diff_ci95_lo",
"sharpe_diff_ci95_hi",
"ret_diff",
"ret_diff_ci95_lo",
"ret_diff_ci95_hi",
"prob_challenger_wins",
"p_value",
}
paired_rows = []
for comparison, frame in (
("holdout minus validation", validation_to_holdout),
("holdout minus equal weight", holdout_to_benchmark),
("validation minus equal weight", validation_to_benchmark),
):
if frame.height != 1 or not paired_required <= set(frame.columns):
raise ValueError(f"missing required paired comparison: {comparison}")
values = frame.row(0, named=True)
if any(values[name] is None or not np.isfinite(values[name]) for name in paired_required):
raise ValueError(f"non-finite paired comparison: {comparison}")
paired_rows.append(
{
"comparison": comparison,
**{name: values[name] for name in paired_required},
}
)
paired_evidence = pl.DataFrame(paired_rows)
paired_evidence
# %% [markdown]
# ## Return and drawdown paths
#
# Each window compounds its own daily returns from initial wealth. The plot does not concatenate the
# disjoint validation and holdout calendars.
# %% tags=["results"]
fig, axes = plt.subplots(1, 2, figsize=(12, 4.2))
for period, result in (("validation", selected_validation), ("holdout", holdout_backtest)):
paths = [path for path in result.artifacts() if path.name == "daily_returns.parquet"]
if len(paths) != 1:
raise ValueError(f"{result.hash} must have one daily return artifact")
returns = pl.read_parquet(paths[0]).select("timestamp", "daily_return").sort("timestamp")
if returns.get_column("timestamp").n_unique() != returns.height:
raise ValueError(f"{result.hash} repeats a return timestamp")
values = returns.get_column("daily_return").to_numpy()
wealth = np.cumprod(1.0 + values)
drawdown = wealth / np.maximum.accumulate(wealth) - 1.0
dates = returns.get_column("timestamp").to_list()
axes[0].plot(dates, wealth, label=period)
axes[1].plot(dates, drawdown, label=period)
axes[0].set_title("Wealth within each evaluation window")
axes[0].set_ylabel("Growth of 1.0")
axes[1].set_title("Drawdown within each evaluation window")
axes[1].set_ylabel("Drawdown")
for axis in axes:
axis.set_xlabel("Date")
axis.legend(frameon=False)
# No `tight_layout` here: matplotlibrc sets `figure.constrained_layout.use`, so the figure
# already has a layout engine and asking for a second one warns. The warning names the
# kernel's own temporary module path, which then lands in the committed notebook and fails
# the output-hygiene guard.
fig.show()
# %% [markdown]
# ## Result interpretation
#
# The sentences below are computed from the registered metrics. An interval wholly above or below zero
# is reported as such; an interval spanning zero is not converted into a positive or negative claim.
# %% tags=["results"]
def _interval_read(name: str, point: float, lower: float, upper: float) -> str:
if lower > 0:
status = "its interval is above zero"
elif upper < 0:
status = "its interval is below zero"
else:
status = "its interval includes zero"
return f"{name}: {point:.3f} [{lower:.3f}, {upper:.3f}]; {status}."
validation_row = selected_performance.filter(pl.col("period") == "validation").row(0, named=True)
holdout_row = selected_performance.filter(pl.col("period") == "holdout").row(0, named=True)
decay_row = paired_evidence.filter(pl.col("comparison") == "holdout minus validation").row(
0, named=True
)
interpretation = pl.DataFrame(
{
"evidence": ["validation", "holdout", "holdout change"],
"reading": [
_interval_read(
"Validation Sharpe",
validation_row["sharpe"],
validation_row["sharpe_ci95_lo"],
validation_row["sharpe_ci95_hi"],
),
_interval_read(
"Holdout Sharpe",
holdout_row["sharpe"],
holdout_row["sharpe_ci95_lo"],
holdout_row["sharpe_ci95_hi"],
),
_interval_read(
"Holdout minus validation Sharpe",
decay_row["sharpe_diff"],
decay_row["sharpe_diff_ci95_lo"],
decay_row["sharpe_diff_ci95_hi"],
),
],
}
)
interpretation
# %% [markdown]
# ## Key takeaways
#
# - The immutable candidate set determines the exact selected configuration, and the holdout
# lineage follows from it rather than being recorded alongside it.
# - Controlled cost and risk evidence changes one strategy field at a time.
# - The holdout cannot cause reselection, because the selection is upstream of it and frozen.
# - All result-specific interpretation is produced from the registered artifacts this notebook
# resolved, so a re-run reports the same numbers or fails.
```स्रोत के लाइसेंस के तहत श्रेय सहित पूरा पाठ दिखाया गया है। लाइसेंस: MIT
यह सारांश मूल स्रोत के आधार पर Stratmill के शोध एजेंट ने लिखा है; यह स्रोत की प्रति नहीं है।