跳至正文
返回文库全部文档

考虑数据泄漏风险的留出集重拟合:FX 预测

代码 《交易机器学习》

总结

本笔记重新拟合已由验证选出的 FX 配置,并登记其预测结果,以供之后的留出集回测使用。它从通过验证且未破产的回测中解析选择结果,并将其限制在冻结的候选集合内,然后依据所选模型不可变的验证规格重建训练规格。留出区间遵循所选标签的观测时间线,训练缓冲区足够宽,可避免标签结果延伸至评估窗口。

笔记只发布选定的检查点,并检查训练标识是否不同于验证拟合、预测是否完整且标记为留出集,以及存储的拟合状态是否通过验证。笔记还报告预测集的覆盖率。这些控制明确了生成的产物,也有助于发现不匹配,但无法阻止反复使用留出集或从多次重跑中挑选结果。文中特别警告,反复运行流程并报告最佳结果会使留出集证据失去可解释性;表现分析留待后续工作。

核心观点

  • 留出集重拟合使用验证选出的配置,而不会重新进行选择。
  • 留出集训练窗口根据标签的观测网格确定,并通过标签缓冲区在评估窗口之前结束。
  • 只发布选定的检查点,避免在留出集检查点之间进行选择。
  • 产物完整性和拟合状态验证分别检查已登记预测的不同属性。
  • 反复运行留出集评估后再报告最佳结果,会削弱评估的意义。

标签

全文
# 17_holdout_predictions.py


```py
# ---
# jupyter:
#   jupytext:
#     cell_metadata_filter: tags,-all
#     text_representation:
#       extension: .py
#       format_name: percent
#       format_version: '1.3'
#       jupytext_version: 1.19.3
#   kernelspec:
#     display_name: Python 3 (ipykernel)
#     language: python
#     name: python3
# ---

# %% [markdown]
# # Holdout Predictions - FX Pairs
#
# This notebook refits the one configuration validation selected, on the holdout interval, and
# registers its predictions. It writes predictions and nothing else: the backtest is
# `18_holdout_backtest`, and what any of it is worth is `19_strategy_analysis`.
#
# Selection is not a parameter and is not made here. `resolve_solvent_carrier` reads the
# highest-Sharpe registered validation backtest across the baseline, allocation and risk-overlay
# stages, restricted to runs that stayed solvent, so this notebook cannot select a configuration
# the validation stages did not already rank first.
#
# The holdout is not a one-shot transaction and nothing here pre-registers it. There is no lock,
# no seal and no gate: the whole rule is retrain the selected configuration on everything up to
# the holdout window, predict, and backtest that same configuration on the result. Re-running is
# therefore ordinary. A reader who runs it five hundred times and quotes the best number has
# produced something uninterpretable, and that is a property of what they did rather than
# something the software can prevent - the earlier design tried to, and bought
# unfixability: a holdout found to be wrong after a bug fix could not be corrected, because the
# lock was by construction the one artifact that could not be revised.
#
# **Learning objectives**
#
# - Derive a holdout interval from the panel's observation grid rather than the calendar.
# - Reconstruct one training identity from an immutable validation specification.
# - Register holdout predictions without giving them any influence over selection.
#
# **Book reference**: Chapters 16-20
#
# **Prerequisite**: `16_costs`. The selected configuration is resolved from the registered
# validation backtests, so every stage that registers one must have run.

# %%
"""Refit the validation-selected FX configuration on the holdout interval."""

import polars as pl

from case_studies.research import open_study
from case_studies.research.comparison import CandidateSet
from case_studies.research.holdout import build_holdout_training_spec
from case_studies.research.models import (
    reconstruct_locked_model_request,
    validate_locked_model_run,
)
from case_studies.utils.strategy_analysis import resolve_solvent_carrier

# %% tags=["parameters"]
CASE_STUDY_ID = "fx_pairs"
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""
CANDIDATE_SET_NAME = "fx_pairs:holdout-candidates"

# %% [markdown]
# ## Resolve the selection and the holdout interval it determines
#
# The label the selection was made on decides which observation grid the holdout interval is
# stepped back along, so it is read from the selected lineage rather than assumed. FX carries
# three labels on one daily grid, which is exactly the coincidence that would let an assumption
# here survive untested.
#
# The training window ends a whole label buffer, counted in observations, before the holdout
# opens, so the last training label's outcome cannot resolve inside the holdout. The buffer is
# the widest this case study configures rather than the primary label's: a 21-day forward return
# resolves three weeks after the session it is stamped on, and a 1-day buffer would leave three
# weeks of holdout outcomes reachable from the training window.

# %% tags=["results"]
study = open_study(CASE_STUDY_ID, execution_tier=EXECUTION_TIER, workspace=WORKSPACE or None)

# The selected configuration is the highest-Sharpe registered VALIDATION backtest across the
# baseline, allocation and risk-overlay stages, restricted to runs that stayed solvent. It is
# resolved from the registry rather than named here, so this notebook cannot select a configuration
# that the validation stages did not rank first. `15_risk_management` froze the set the holdout is
# allowed to choose from, and that set is passed into the resolution rather than checked against its
# answer. The two are different tests: when a conformal candidate is in the field the resolver
# re-ranks every candidate on the timestamps they all share, so a row that was never admitted still
# decides how far that intersection reaches and therefore which admitted row wins. fx has 194
# conformal backtests, so this is a live path here rather than a hypothetical one. Checking
# membership afterwards would pass while the answer had already been changed by an ineligible row.
holdout_candidates = CandidateSet.one(study, name=CANDIDATE_SET_NAME)
carrier = resolve_solvent_carrier(CASE_STUDY_ID, admitted=frozenset(holdout_candidates.members))
print(
    f"Frozen candidate set {holdout_candidates.hash}: "
    f"{len(holdout_candidates.members)} members, "
    f"raw-Sharpe pick {holdout_candidates.best_validation_sharpe().hash}"
)

validation_prediction = study.results.open(carrier["val_prediction_hash"])
prediction_record = validation_prediction.registry_record()
CHECKPOINT_KIND = prediction_record["checkpoint_kind"]
CHECKPOINT_VALUE = prediction_record["checkpoint_value"]

# The label the selected configuration was fitted on decides which observation grid the holdout
# interval is stepped back along, so it is read from the selected configuration rather than assumed.
# FX carries three labels on one daily grid, which is exactly the coincidence that would let an
# assumption here survive untested.
observation_timeline = (
    pl.read_parquet(study.root / "labels" / f"{carrier['label']}.parquet")
    .get_column("timestamp")
    .unique()
    .sort()
    .to_list()
)
validation_spec = study.results.open(carrier["training_hash"]).spec()
holdout_spec = build_holdout_training_spec(
    study,
    validation_spec,
    timeline=observation_timeline,
    case_study=CASE_STUDY_ID,
)
holdout_fold = holdout_spec["computation"]["cv"]["folds"][0]

pl.DataFrame(
    {
        "field": [
            "selected backtest",
            "selected stage",
            "validation Sharpe",
            "family",
            "configuration",
            "label",
            "checkpoint",
            "validation training",
            "validation prediction",
            "holdout train window",
            "holdout evaluation window",
        ],
        "value": [
            carrier["val_backtest_hash"],
            str(carrier["val_stage"]),
            f"{carrier['val_sharpe']:.4f}",
            str(carrier["family"]),
            str(carrier["config_name"]),
            str(carrier["label"]),
            f"{CHECKPOINT_KIND}={CHECKPOINT_VALUE}",
            str(carrier["training_hash"]),
            validation_prediction.hash,
            f"{holdout_fold['train_start']} to {holdout_fold['train_end']}",
            f"{holdout_fold['val_start']} to {holdout_fold['val_end']}",
        ],
    }
)

# %% [markdown]
# ## Refit the selected configuration on the holdout fold
#
# The request is reconstructed from the immutable validation specification with only the fold
# geometry re-keyed, so the holdout model differs from the validation model in what it was
# fitted on and in nothing else. It publishes the selected checkpoint alone: a holdout refit
# that published its whole checkpoint schedule would hand the next notebook a choice, and
# choosing among holdout checkpoints is selection on the holdout under another name.
#
# The one thing checked afterwards that a specification cannot state about itself is that this
# is a refit at all. A holdout training identity equal to the validation one means the fold
# re-keying changed nothing, and the model is a validation fit predicting forward over a later
# window rather than a model trained up to it.

# %% tags=["results"]
request = reconstruct_locked_model_request(
    study,
    holdout_spec,
    checkpoint_kind=CHECKPOINT_KIND,
    checkpoint_value=CHECKPOINT_VALUE,
)
model_run = request.run()
if model_run.training.hash == carrier["training_hash"]:
    raise RuntimeError(
        f"the holdout refit produced the validation training identity "
        f"{carrier['training_hash']}, so it did not refit"
    )
if len(model_run.predictions) != 1:
    raise RuntimeError(
        f"the holdout refit published {len(model_run.predictions)} prediction sets; "
        "only the selected checkpoint may be published"
    )
prediction = model_run.predictions[0]

record = prediction.registry_record()
if record["split"] != "holdout":
    raise RuntimeError(f"the holdout refit published a {record['split']!r} prediction")
if record["checkpoint_kind"] != CHECKPOINT_KIND or record["checkpoint_value"] != CHECKPOINT_VALUE:
    raise RuntimeError(
        f"the holdout prediction is at checkpoint {record['checkpoint_kind']}="
        f"{record['checkpoint_value']}, not the carrier's {CHECKPOINT_KIND}={CHECKPOINT_VALUE}"
    )
if not prediction.complete:
    raise RuntimeError("the holdout prediction is incomplete")

# Completeness verifies the published prediction artifact, not that the persisted fitted state
# reproduces it. A cached run whose model state is missing or inconsistent passes the check
# above and fails only where someone tries to use the model again. The family's own validator
# reads that state and returns its digest, which is the one thing about this run that the
# specification cannot state about itself.
fitted_state_digest = validate_locked_model_run(request, model_run)

print(f"Holdout training run:   {model_run.training.hash}")
print(f"Fitted-state digest:    {fitted_state_digest}")
print(f"Holdout prediction set: {prediction.hash}")

# %% [markdown]
# ## What the holdout refit covers
#
# The coverage is printed rather than assumed: a holdout prediction set that silently covers a
# shorter window than the fold declares would make every number downstream a measurement of a
# different interval than the one this notebook says it measured.

# %% tags=["results"]
frame = prediction.load()
coverage = pl.DataFrame(
    {
        "field": ["prediction", "rows", "symbols", "sessions", "first session", "last session"],
        "value": [
            prediction.hash,
            str(frame.height),
            str(frame.get_column("symbol").n_unique()),
            str(frame.get_column("timestamp").n_unique()),
            str(frame.get_column("timestamp").min()),
            str(frame.get_column("timestamp").max()),
        ],
    }
)
coverage

# %% [markdown]
# ## Key takeaways
#
# - The configuration refitted here was selected on validation, from an immutable set, upstream.
# - The holdout training window stops a full label buffer short of the holdout, counted in
#   observations rather than calendar days.
# - Only the selected checkpoint is published, so no choice among holdout results remains to be
#   made downstream.

```

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: MIT

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。