跳至正文
返回文库全部文档

基于验证结果选定 FX 后重拟合留出集,不用留出集选模

笔记本 《交易机器学习》

总结

本笔记说明如何在通过验证结果选定 FX 策略配置后生成留出集预测。它会在获准且保持偿付能力的候选项中,选定验证回测夏普比率最高的一项,然后根据所选标签的观测时间线构建留出区间。训练会在留出集之前预留一段标签缓冲区,避免训练结果延伸至评估期。笔记根据验证规格重建模型请求,在保留所选配置和检查点的同时调整折的几何结构。

本笔记检查重拟合是否具有独立的训练身份、是否仅发布选定的检查点,以及是否生成覆盖范围符合预期的完整留出集预测。它还会验证已持久化的拟合模型状态。这些保护措施可确定拟合内容和预测区间,但无法防止重复使用留出集或选择性报告。反复运行留出集并挑选最佳结果会损害其可解释性,因此后续表现分析另行处理。

核心观点

  • 生成留出集预测前,先依据验证结果选定 FX 配置。
  • 根据所选标签的观测网格确定留出集边界,并在评估前留出完整的标签缓冲区。
  • 根据验证规格重建留出集拟合,只调整折的几何结构。
  • 仅发布选定的检查点,避免留出集结果带来额外的模型选择。
  • 反复测试并选择性报告留出集运行结果,会使评估难以解读。

标签

全文
# Holdout Predictions - FX Pairs


# Holdout Predictions - FX Pairs

This notebook refits the one configuration validation selected, on the holdout interval, and
registers its predictions. It writes predictions and nothing else: the backtest is
`18_holdout_backtest`, and what any of it is worth is `19_strategy_analysis`.

Selection is not a parameter and is not made here. `resolve_solvent_carrier` reads the
highest-Sharpe registered validation backtest across the baseline, allocation and risk-overlay
stages, restricted to runs that stayed solvent, so this notebook cannot select a configuration
the validation stages did not already rank first.

The holdout is not a one-shot transaction and nothing here pre-registers it. There is no lock,
no seal and no gate: the whole rule is retrain the selected configuration on everything up to
the holdout window, predict, and backtest that same configuration on the result. Re-running is
therefore ordinary. A reader who runs it five hundred times and quotes the best number has
produced something uninterpretable, and that is a property of what they did rather than
something the software can prevent - the earlier design tried to, and bought
unfixability: a holdout found to be wrong after a bug fix could not be corrected, because the
lock was by construction the one artifact that could not be revised.

**Learning objectives**

- Derive a holdout interval from the panel's observation grid rather than the calendar.
- Reconstruct one training identity from an immutable validation specification.
- Register holdout predictions without giving them any influence over selection.

**Book reference**: Chapters 16-20

**Prerequisite**: `16_costs`. The selected configuration is resolved from the registered
validation backtests, so every stage that registers one must have run.

```python
"""Refit the validation-selected FX configuration on the holdout interval."""

import polars as pl

from case_studies.research import open_study
from case_studies.research.comparison import CandidateSet
from case_studies.research.holdout import build_holdout_training_spec
from case_studies.research.models import (
    reconstruct_locked_model_request,
    validate_locked_model_run,
)
from case_studies.utils.strategy_analysis import resolve_solvent_carrier
```

```python
CASE_STUDY_ID = "fx_pairs"
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""
CANDIDATE_SET_NAME = "fx_pairs:holdout-candidates"
```

## Resolve the selection and the holdout interval it determines

The label the selection was made on decides which observation grid the holdout interval is
stepped back along, so it is read from the selected lineage rather than assumed. FX carries
three labels on one daily grid, which is exactly the coincidence that would let an assumption
here survive untested.

The training window ends a whole label buffer, counted in observations, before the holdout
opens, so the last training label's outcome cannot resolve inside the holdout. The buffer is
the widest this case study configures rather than the primary label's: a 21-day forward return
resolves three weeks after the session it is stamped on, and a 1-day buffer would leave three
weeks of holdout outcomes reachable from the training window.

```python
study = open_study(CASE_STUDY_ID, execution_tier=EXECUTION_TIER, workspace=WORKSPACE or None)

# The selected configuration is the highest-Sharpe registered VALIDATION backtest across the
# baseline, allocation and risk-overlay stages, restricted to runs that stayed solvent. It is
# resolved from the registry rather than named here, so this notebook cannot select a configuration
# that the validation stages did not rank first. `15_risk_management` froze the set the holdout is
# allowed to choose from, and that set is passed into the resolution rather than checked against its
# answer. The two are different tests: when a conformal candidate is in the field the resolver
# re-ranks every candidate on the timestamps they all share, so a row that was never admitted still
# decides how far that intersection reaches and therefore which admitted row wins. fx has 194
# conformal backtests, so this is a live path here rather than a hypothetical one. Checking
# membership afterwards would pass while the answer had already been changed by an ineligible row.
holdout_candidates = CandidateSet.one(study, name=CANDIDATE_SET_NAME)
carrier = resolve_solvent_carrier(CASE_STUDY_ID, admitted=frozenset(holdout_candidates.members))
print(
    f"Frozen candidate set {holdout_candidates.hash}: "
    f"{len(holdout_candidates.members)} members, "
    f"raw-Sharpe pick {holdout_candidates.best_validation_sharpe().hash}"
)

validation_prediction = study.results.open(carrier["val_prediction_hash"])
prediction_record = validation_prediction.registry_record()
CHECKPOINT_KIND = prediction_record["checkpoint_kind"]
CHECKPOINT_VALUE = prediction_record["checkpoint_value"]

# The label the selected configuration was fitted on decides which observation grid the holdout
# interval is stepped back along, so it is read from the selected configuration rather than assumed.
# FX carries three labels on one daily grid, which is exactly the coincidence that would let an
# assumption here survive untested.
observation_timeline = (
    pl.read_parquet(study.root / "labels" / f"{carrier['label']}.parquet")
    .get_column("timestamp")
    .unique()
    .sort()
    .to_list()
)
validation_spec = study.results.open(carrier["training_hash"]).spec()
holdout_spec = build_holdout_training_spec(
    study,
    validation_spec,
    timeline=observation_timeline,
    case_study=CASE_STUDY_ID,
)
holdout_fold = holdout_spec["computation"]["cv"]["folds"][0]

pl.DataFrame(
    {
        "field": [
            "selected backtest",
            "selected stage",
            "validation Sharpe",
            "family",
            "configuration",
            "label",
            "checkpoint",
            "validation training",
            "validation prediction",
            "holdout train window",
            "holdout evaluation window",
        ],
        "value": [
            carrier["val_backtest_hash"],
            str(carrier["val_stage"]),
            f"{carrier['val_sharpe']:.4f}",
            str(carrier["family"]),
            str(carrier["config_name"]),
            str(carrier["label"]),
            f"{CHECKPOINT_KIND}={CHECKPOINT_VALUE}",
            str(carrier["training_hash"]),
            validation_prediction.hash,
            f"{holdout_fold['train_start']} to {holdout_fold['train_end']}",
            f"{holdout_fold['val_start']} to {holdout_fold['val_end']}",
        ],
    }
)
```

## Refit the selected configuration on the holdout fold

The request is reconstructed from the immutable validation specification with only the fold
geometry re-keyed, so the holdout model differs from the validation model in what it was
fitted on and in nothing else. It publishes the selected checkpoint alone: a holdout refit
that published its whole checkpoint schedule would hand the next notebook a choice, and
choosing among holdout checkpoints is selection on the holdout under another name.

The one thing checked afterwards that a specification cannot state about itself is that this
is a refit at all. A holdout training identity equal to the validation one means the fold
re-keying changed nothing, and the model is a validation fit predicting forward over a later
window rather than a model trained up to it.

```python
request = reconstruct_locked_model_request(
    study,
    holdout_spec,
    checkpoint_kind=CHECKPOINT_KIND,
    checkpoint_value=CHECKPOINT_VALUE,
)
model_run = request.run()
if model_run.training.hash == carrier["training_hash"]:
    raise RuntimeError(
        f"the holdout refit produced the validation training identity "
        f"{carrier['training_hash']}, so it did not refit"
    )
if len(model_run.predictions) != 1:
    raise RuntimeError(
        f"the holdout refit published {len(model_run.predictions)} prediction sets; "
        "only the selected checkpoint may be published"
    )
prediction = model_run.predictions[0]

record = prediction.registry_record()
if record["split"] != "holdout":
    raise RuntimeError(f"the holdout refit published a {record['split']!r} prediction")
if record["checkpoint_kind"] != CHECKPOINT_KIND or record["checkpoint_value"] != CHECKPOINT_VALUE:
    raise RuntimeError(
        f"the holdout prediction is at checkpoint {record['checkpoint_kind']}="
        f"{record['checkpoint_value']}, not the carrier's {CHECKPOINT_KIND}={CHECKPOINT_VALUE}"
    )
if not prediction.complete:
    raise RuntimeError("the holdout prediction is incomplete")

# Completeness verifies the published prediction artifact, not that the persisted fitted state
# reproduces it. A cached run whose model state is missing or inconsistent passes the check
# above and fails only where someone tries to use the model again. The family's own validator
# reads that state and returns its digest, which is the one thing about this run that the
# specification cannot state about itself.
fitted_state_digest = validate_locked_model_run(request, model_run)

print(f"Holdout training run:   {model_run.training.hash}")
print(f"Fitted-state digest:    {fitted_state_digest}")
print(f"Holdout prediction set: {prediction.hash}")
```

## What the holdout refit covers

The coverage is printed rather than assumed: a holdout prediction set that silently covers a
shorter window than the fold declares would make every number downstream a measurement of a
different interval than the one this notebook says it measured.

```python
frame = prediction.load()
coverage = pl.DataFrame(
    {
        "field": ["prediction", "rows", "symbols", "sessions", "first session", "last session"],
        "value": [
            prediction.hash,
            str(frame.height),
            str(frame.get_column("symbol").n_unique()),
            str(frame.get_column("timestamp").n_unique()),
            str(frame.get_column("timestamp").min()),
            str(frame.get_column("timestamp").max()),
        ],
    }
)
coverage
```

## Key takeaways

- The configuration refitted here was selected on validation, from an immutable set, upstream.
- The holdout training window stops a full label buffer short of the holdout, counted in
  observations rather than calendar days.
- Only the selected checkpoint is published, so no choice among holdout results remains to be
  made downstream.

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: MIT

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。