跳至正文
返回文库全部文档

留出集预测前重新拟合选定模型

代码 《交易机器学习》

总结

本笔记介绍一项US公司特征研究中的留出集预测阶段。流程采用根据验证结果选定的模型配置,重建其训练规格和检查点,然后使用截至留出窗口之前的数据重新拟合。新的训练标识使重新拟合过程可审计;窗口和标签缓冲区则根据研究设置推导,以降低训练结果与评估期重叠的风险。

笔记还处理注册表完整性问题:找出一个复用验证训练标识的早期预测版本,并说明如何替换版本,使下游用户不必在相互竞争的结果中选择。笔记确认已根据选定配置生成预测,但没有为其评分或据此交易;这些步骤属于后续分析。该配置已经通过验证结果选定,因此留出期并非完全未触及的研究过程。此处描述的是一个案例研究,不能证明预测质量或策略盈利能力。

核心观点

  • 在生成留出集预测前,使用验证结果选定配置。
  • 使用截至留出期之前的数据重新拟合模型,并设置标签缓冲区以避免结果重叠。
  • 独立的训练标识能够证明执行了新的留出集拟合。
  • 笔记生成预测,但不评估其表现,也不将其转换为收益。
  • 先前基于验证结果进行选择,限制了将留出集严格描述为未触及数据的程度。

标签

全文
# 15_holdout_predictions.py


```py
# ---
# jupyter:
#   jupytext:
#     cell_metadata_filter: tags,-all
#     text_representation:
#       extension: .py
#       format_name: percent
#       format_version: '1.3'
#       jupytext_version: 1.19.3
#   kernelspec:
#     display_name: Python 3 (ipykernel)
#     language: python
#     name: python3
# ---

# %% [markdown]
# # US Firm Characteristics: Holdout Predictions
#
# **Chapter 20 - Out-of-sample evaluation**
#
# Every number in this case study so far was measured on the validation folds, and every
# choice was made by looking at them: which model family, which configuration, how many
# names to hold, how to size them. A result selected that way cannot also be evidence that
# the selection was sound - the ranking and the evidence would be the same measurement.
#
# The holdout is the window nothing has been selected on. This notebook fits the selected
# configuration on the history available before that window opens and writes its
# predictions over it. [`16_holdout_backtest`](16_holdout_backtest.ipynb) turns those
# predictions into a return series with the sizing the case study settled on, and
# [`17_strategy_analysis`](17_strategy_analysis.ipynb) reads both back.
#
# **What this notebook is careful about**
#
# A holdout prediction is not the validation model scored on a later window. That is the
# mistake this case study had already made: the registry carried a holdout prediction set
# generated from the same training identity as the validation run, so what it scored was a
# model whose parameters had been chosen while looking at the folds it was being judged
# against. Section 2 fits again, and the new training identity is what makes the refit
# visible rather than asserted.
#
# **Prerequisites:** [`14_costs`](14_costs.ipynb), which fixes the configuration the
# holdout runs.
#
# **Scope:** one training run and one prediction set. No backtest, no selection, no
# comparison - those are 16 and 17.

# %%
"""US Firm Characteristics: Holdout Predictions."""

import polars as pl

from case_studies.research import open_study
from case_studies.research.holdout import build_holdout_training_spec
from case_studies.research.models import reconstruct_locked_model_request
from case_studies.utils.registry import training_hash_from_spec
from case_studies.utils.registry.maintenance import delete_prediction_generation
from case_studies.utils.strategy_analysis import (
    holdout_generations_to_retire,
    registered_holdout_generations,
    resolve_solvent_carrier,
)
from case_studies.utils.warning_policy import apply_notebook_warning_policy
from utils.paths import get_case_study_dir

apply_notebook_warning_policy()

# %% tags=["parameters"]
CASE_STUDY_ID = "us_firm_characteristics"
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""
# Whether a holdout generation for a DIFFERENT configuration may be superseded by this run.
# Off by default: see section 3.
#
# The flag exists because the holdout is not a one-shot resource, which is a ruling and not
# an oversight. What the rule against consulting the holdout forbids is SELECTING on it: the
# configuration evaluated here is chosen by validation backtest Sharpe, and no holdout number
# feeds back into that choice. It says nothing about how many times the evaluation may be
# computed, and a wrong result is deleted and re-run rather than left standing because it was
# observed. Reading the rule as a physical constraint is what produced a lock layer around
# this window, and it is being removed. The guard here is against something narrower and real:
# two generations readable at once, so nobody downstream has to choose between them and nobody
# can quote whichever number they prefer.
REPLACE_HOLDOUT = False

# %%
study = open_study(CASE_STUDY_ID, execution_tier=EXECUTION_TIER, workspace=WORKSPACE or None)
CASE_DIR = get_case_study_dir(CASE_STUDY_ID)


def _delete_holdout_generation(case_dir, prediction_hash):
    """Remove one holdout prediction set and everything registered against it.

    The rows go rather than being marked, because a holdout evaluation that is still
    readable is still a number someone can quote, and the point of removing it is that it
    should not be one.

    The cascade lives in `case_studies/utils/registry/maintenance.py` and derives the child
    tables from the schema. The version this replaces listed them by hand and was already
    missing `cohort_metrics.leader_hash`, which with foreign keys enabled aborts the delete
    rather than orphaning a row - so on any registry with a cohort this function raised
    instead of deleting, and the generation it was called on stayed.
    """
    deleted = delete_prediction_generation(case_dir / "run_log" / "registry.db", prediction_hash)
    for table, n in sorted(deleted.items()):
        print(f"  deleted {n:>3} from {table}")


# %% [markdown]
# ## 1. Which configuration the holdout runs
#
# The holdout runs the configuration the case study reports, resolved through the same
# `resolve_solvent_carrier` [`14_costs`](14_costs.ipynb) uses. Resolving it again here
# rather than passing it along is deliberate: the two notebooks must agree by construction,
# and a hash written down in one and read in the other agrees only until the sweep is
# rebuilt.
#
# Nothing about the holdout enters this choice. The selected configuration is the validation
# rank-1, and it was fixed before this notebook ran.

# %%
carrier = resolve_solvent_carrier(CASE_STUDY_ID)
print(
    f"Selected configuration: {carrier['val_backtest_hash']}  stage={carrier['val_stage']}  "
    f"family={carrier['family']}  config={carrier['config_name']}  "
    f"label={carrier['label']}"
)
print(
    f"  validation Sharpe {carrier['val_sharpe']:.3f}, max drawdown {carrier['max_drawdown']:.3f}"
)
print(f"  fitted by training run {carrier['training_hash']}")

# %% [markdown]
# The checkpoint is part of the configuration. This family publishes a prediction set per boosting
# iteration on a declared schedule, and the selected configuration's prediction set names one of
# them - so refitting without it would produce a model at the end of training rather than the one
# that was ranked.

# %%
validation_prediction = study.results.open(carrier["val_prediction_hash"])
prediction_record = validation_prediction.registry_record()
CHECKPOINT_KIND = prediction_record["checkpoint_kind"]
CHECKPOINT_VALUE = prediction_record["checkpoint_value"]
print(f"Checkpoint: {CHECKPOINT_KIND}={CHECKPOINT_VALUE}")

# %% [markdown]
# ## 2. The window, and the model that is allowed to see it
#
# The holdout window is not a choice made here. It is `evaluation.holdout_start` and
# `evaluation.holdout_end` from the case study's own `setup.yaml`, read through the same
# `canonical_window` the fold derivation and the backtest slice both go through, so the
# three cannot disagree.
#
# The training interval is everything available before that window, bounded above by a
# label buffer. The buffer is what stops the last training label's outcome from resolving
# inside the holdout: this case study dates each row by the month the return was earned,
# so a monthly label observed at the end of December is already realised, and the buffer
# is one observation rather than a horizon's worth. A zero gap would be a leak, not a
# conservative choice, so the derivation refuses to default it.
#
# Everything else about the configuration is carried across unchanged, and the fields that
# cannot be - the eligibility manifest, and any parameter this family resolves from a
# fold's own training rows - are recomputed against the holdout fold. Carrying those
# forward would fit a model keyed to the validation folds and call it a retrain.

# %%
observation_timeline = (
    pl.read_parquet(study.root / "labels" / f"{carrier['label']}.parquet")
    .get_column("timestamp")
    .unique()
    .sort()
    .to_list()
)
validation_spec = study.results.open(carrier["training_hash"]).spec()
holdout_spec = build_holdout_training_spec(
    study,
    validation_spec,
    timeline=observation_timeline,
    case_study=CASE_STUDY_ID,
)

fold = holdout_spec["computation"]["cv"]["folds"][0]
print(f"Holdout fold {fold['fold']}")
print(f"  trains  {fold['train_start']} -> {fold['train_end']}")
print(f"  predicts {fold['val_start']} -> {fold['val_end']}")
print(f"  label buffer: {holdout_spec['computation']['cv']['request']['label_buffer']}")

# The validation folds are what the buffer is measured against, and the last of them ends
# before the holdout opens. Printing both is what lets a reader check the gap rather than
# take it on the derivation's word.
validation_folds = validation_spec["computation"]["cv"]["folds"]
latest_validation_end = max(str(entry["val_end"]) for entry in validation_folds)
print(f"Validation folds: {len(validation_folds)}, latest evaluation end {latest_validation_end}")
print(f"Holdout training ends {fold['train_end']}, holdout opens {fold['val_start']}")

# %% [markdown]
# ## 3. Fit, and register the predictions
#
# `reconstruct_locked_model_request` builds the request from the spec above. Its name
# comes from the locked holdout path this case study no longer uses; it takes a training
# specification and a checkpoint, not a lock, and it is used here because it is the one
# call that refuses a request that is not exactly the spec it was handed - the training
# identity, the checkpoint schedule, the feature lineage and the runtime parameters are
# all checked before anything is fitted.
#
# The training identity below is new. It has to be: it covers the CV interval, and the
# holdout fold is not one of the validation folds. A run that came back with the
# validation training hash would mean the refit did not happen.
#
# **The window carries one configuration at a time.** The holdout is re-runnable, and that
# is not the same as free: every configuration evaluated on it is another look at a period
# the case study reports as unseen, and two evaluated quietly would make that report false.
#
# So the check below is on the selected configuration rather than on the notebook, and it has
# exactly two outcomes. With the selected configuration unchanged this is an idempotent replay: the
# derivation is deterministic and the training identity covers it, so the same identity comes back
# and the fit is served from the registry. With the selected configuration changed it refuses,
# names both configurations, and stops.
#
# `REPLACE_HOLDOUT` is the only way past that, and it is a replacement rather than an
# addition: the superseded generation's rows are deleted, so the registry never holds two
# refits of the holdout window and no downstream resolver has to choose between them.
# Deleting is what makes the earlier evaluation cost something to discard. It is also the
# only honest shape - a run that had been observed and then quietly kept alongside its
# replacement would let a reader take whichever number they preferred.

# %%
holdout_training_hash = training_hash_from_spec(holdout_spec)
this_generation = (holdout_training_hash, (CHECKPOINT_KIND, CHECKPOINT_VALUE))
retire = holdout_generations_to_retire(CASE_DIR, this_generation=this_generation)
# A row whose training run records no CV split cannot be shown either way, and deleting on
# that would discard a result nothing has established is wrong. It stops the run instead.
if retire.unattributable:
    raise RuntimeError(
        "the holdout window carries prediction sets whose training runs record no CV split, "
        "so whether they were refitted for the holdout cannot be established: "
        + ", ".join(
            f"{row['prediction_hash']} (training {row['training_hash']})"
            for row in retire.unattributable
        )
        + ". Establish what produced them before registering another evaluation on the same "
        "window; this notebook will not delete a row it cannot show is not a holdout result."
    )
# A row whose training run declares a non-holdout CV may not be reported as a holdout
# result, and it is also not something to delete unattended: `generate_holdout` refits on a
# holdout fold and then registers the predictions under the VALIDATION training identity, so
# this record covers both a validation-fitted model published over the window and a real
# refit filed under the wrong identity. Nothing owned this before - the filter here was
# `row["refitted"]`, which made exactly these rows invisible to the refusal and to
# everything after it.
if retire.not_out_of_sample and not REPLACE_HOLDOUT:
    raise RuntimeError(
        "the holdout window carries prediction sets whose training runs declare a CV split "
        "other than the holdout: "
        + ", ".join(
            f"{row['prediction_hash']} ({row['config_name']}, training {row['training_hash']})"
            for row in retire.not_out_of_sample
        )
        + ". Each is either a validation-fitted model published over the window, which is "
        "not an out-of-sample result, or a refit registered under its validation training "
        "identity, which `20_strategy_synthesis/holdout.py::generate_holdout` produces - and "
        "the registry cannot tell those apart. Establish which, then set REPLACE_HOLDOUT="
        "True to remove it, or leave it and resolve the identity instead."
    )
for row in retire.not_out_of_sample:
    print(
        f"REMOVING {row['prediction_hash']} ({row['config_name']}, training "
        f"{row['training_hash']}): its training run declares a non-holdout CV, so it is not "
        "reportable as a holdout evaluation under the identity it carries"
    )
    _delete_holdout_generation(CASE_DIR, row["prediction_hash"])
superseded = list(retire.superseded)
if superseded and not REPLACE_HOLDOUT:
    raise RuntimeError(
        "the holdout window already carries a refit of a different configuration: "
        + ", ".join(
            f"{row['prediction_hash']} ({row['config_name']}, training {row['training_hash']})"
            for row in superseded
        )
        + f". This run would evaluate {carrier['config_name']} (training "
        f"{holdout_training_hash}, checkpoint {CHECKPOINT_KIND}={CHECKPOINT_VALUE}) on the "
        "same window. Set REPLACE_HOLDOUT=True to discard the earlier generation, or leave "
        "the selection where it was."
    )
for row in superseded:
    print(f"REPLACING holdout generation {row['prediction_hash']} ({row['config_name']})")
    _delete_holdout_generation(CASE_DIR, row["prediction_hash"])

# %% tags=["results"]
request = reconstruct_locked_model_request(
    study,
    holdout_spec,
    checkpoint_kind=CHECKPOINT_KIND,
    checkpoint_value=CHECKPOINT_VALUE,
)
model_run = request.run()
holdout_prediction = model_run.predictions[0]

if model_run.training.hash == carrier["training_hash"]:
    raise RuntimeError(
        "the holdout refit produced the validation training identity "
        f"{carrier['training_hash']}, which means it did not refit"
    )
print(f"Holdout training run:  {model_run.training.hash}")
print(f"Holdout prediction set: {holdout_prediction.hash}")

# %% [markdown]
# What the prediction set covers, read back from the registry rather than from the
# request. The two agree only if the fit published what it declared, and the count is the
# one number a reader can check the window against: a monthly panel over one year is
# twelve decision dates, and the number of rows is those dates times the names eligible on
# each.

# %% tags=["results"]
record = holdout_prediction.registry_record()
predictions = holdout_prediction.load()
print(
    f"split={record['split']}  checkpoint={record['checkpoint_kind']}={record['checkpoint_value']}"
)
print(f"rows={predictions.height:,}  dates={predictions['timestamp'].n_unique()}")
print(
    f"  {predictions['timestamp'].min()} -> {predictions['timestamp'].max()}, "
    f"{predictions['symbol'].n_unique():,} names"
)

# %% [markdown]
# The registry now holds more than one holdout prediction set for this case study, and
# only one of them was fitted on data that ends before the window. The other is the
# defective generation this notebook replaces: it carries the validation training identity,
# which is how it was found. Both are listed rather than one silently preferred, because
# the registry is immutable and a reader looking at it later will see both.

# %% tags=["results"]
for row in registered_holdout_generations(CASE_DIR):
    note = (
        "refitted for the holdout" if row["refitted"] else "VALIDATION-FITTED - not out of sample"
    )
    print(
        f"  {row['prediction_hash']}  training={row['training_hash']}  {row['config_name']}  {note}"
    )

# %% [markdown]
# ## What this notebook establishes, and what it does not
#
# It establishes one thing: a prediction set over the holdout window, produced by the
# configuration this case study selected, fitted on data that ends before the window
# opens. That is a precondition for an out-of-sample claim, not the claim itself. Nothing
# here says whether the predictions are any good - they have not been scored, sized or
# traded.
#
# It does not make the holdout a fresh test in the strict sense. The configuration reached
# this notebook through a selection made on the validation folds, and this window is being
# used once per configuration that gets here. What it does remove is the specific
# circularity of scoring a validation-fitted model on the period meant to judge it.
#
# The holdout is re-runnable. If a later pass finds the selection was wrong, the answer is
# to delete this generation and produce another, not to treat the first as spent.
#
# **Next:** [`16_holdout_backtest`](16_holdout_backtest.ipynb).

```

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: MIT

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。