본문으로 건너뛰기
라이브러리 문서 전체

홀드아웃 예측 전 선택된 모델 재적합

코드 Machine Learning for Trading

요약

이 노트북은 US 기업 특성 연구의 홀드아웃 예측 단계를 설명합니다. 검증 결과로 선택한 모델 설정의 학습 사양과 체크포인트를 복원한 뒤, 홀드아웃 기간 전에 끝나는 데이터로 다시 적합합니다. 새 학습 식별자를 사용해 재적합을 감사할 수 있게 하고, 연구 설정에서 기간과 레이블 버퍼를 정해 학습 결과와 평가 기간이 겹칠 위험을 줄입니다.

노트북은 레지스트리 무결성도 다룹니다. 검증 학습 식별자를 재사용한 이전 예측 생성을 찾아내고, 후속 사용자가 경쟁 결과 중 하나를 고르지 않아도 되도록 생성 결과를 교체하는 방법을 설명합니다. 선택한 설정으로 예측을 생성했다는 점은 확인하지만, 점수를 매기거나 이를 거래에 사용하지는 않습니다. 그런 단계는 이후 분석에서 다룹니다. 설정은 이미 검증 과정을 거쳐 선택되었으므로 홀드아웃을 완전히 손대지 않은 연구 과정이라고 할 수는 없습니다. 설명은 하나의 사례 연구에 관한 것이며 예측 품질이나 전략 수익성을 입증하지 않습니다.

핵심 아이디어

  • 홀드아웃 예측을 생성하기 전에 검증 결과를 사용해 설정을 선택합니다.
  • 결과가 겹치지 않도록 레이블 버퍼를 두고 홀드아웃 전에 끝나는 데이터로 모델을 다시 적합합니다.
  • 별도의 학습 식별자는 새로운 홀드아웃 적합을 수행했다는 근거가 됩니다.
  • 노트북은 예측을 생성하지만 성능을 평가하거나 수익률로 전환하지는 않습니다.
  • 검증 기반 사전 선택이 있었으므로 홀드아웃을 완전히 손대지 않았다고 엄격히 표현하기는 어렵습니다.

태그

전문
# 15_holdout_predictions.py


```py
# ---
# jupyter:
#   jupytext:
#     cell_metadata_filter: tags,-all
#     text_representation:
#       extension: .py
#       format_name: percent
#       format_version: '1.3'
#       jupytext_version: 1.19.3
#   kernelspec:
#     display_name: Python 3 (ipykernel)
#     language: python
#     name: python3
# ---

# %% [markdown]
# # US Firm Characteristics: Holdout Predictions
#
# **Chapter 20 - Out-of-sample evaluation**
#
# Every number in this case study so far was measured on the validation folds, and every
# choice was made by looking at them: which model family, which configuration, how many
# names to hold, how to size them. A result selected that way cannot also be evidence that
# the selection was sound - the ranking and the evidence would be the same measurement.
#
# The holdout is the window nothing has been selected on. This notebook fits the selected
# configuration on the history available before that window opens and writes its
# predictions over it. [`16_holdout_backtest`](16_holdout_backtest.ipynb) turns those
# predictions into a return series with the sizing the case study settled on, and
# [`17_strategy_analysis`](17_strategy_analysis.ipynb) reads both back.
#
# **What this notebook is careful about**
#
# A holdout prediction is not the validation model scored on a later window. That is the
# mistake this case study had already made: the registry carried a holdout prediction set
# generated from the same training identity as the validation run, so what it scored was a
# model whose parameters had been chosen while looking at the folds it was being judged
# against. Section 2 fits again, and the new training identity is what makes the refit
# visible rather than asserted.
#
# **Prerequisites:** [`14_costs`](14_costs.ipynb), which fixes the configuration the
# holdout runs.
#
# **Scope:** one training run and one prediction set. No backtest, no selection, no
# comparison - those are 16 and 17.

# %%
"""US Firm Characteristics: Holdout Predictions."""

import polars as pl

from case_studies.research import open_study
from case_studies.research.holdout import build_holdout_training_spec
from case_studies.research.models import reconstruct_locked_model_request
from case_studies.utils.registry import training_hash_from_spec
from case_studies.utils.registry.maintenance import delete_prediction_generation
from case_studies.utils.strategy_analysis import (
    holdout_generations_to_retire,
    registered_holdout_generations,
    resolve_solvent_carrier,
)
from case_studies.utils.warning_policy import apply_notebook_warning_policy
from utils.paths import get_case_study_dir

apply_notebook_warning_policy()

# %% tags=["parameters"]
CASE_STUDY_ID = "us_firm_characteristics"
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""
# Whether a holdout generation for a DIFFERENT configuration may be superseded by this run.
# Off by default: see section 3.
#
# The flag exists because the holdout is not a one-shot resource, which is a ruling and not
# an oversight. What the rule against consulting the holdout forbids is SELECTING on it: the
# configuration evaluated here is chosen by validation backtest Sharpe, and no holdout number
# feeds back into that choice. It says nothing about how many times the evaluation may be
# computed, and a wrong result is deleted and re-run rather than left standing because it was
# observed. Reading the rule as a physical constraint is what produced a lock layer around
# this window, and it is being removed. The guard here is against something narrower and real:
# two generations readable at once, so nobody downstream has to choose between them and nobody
# can quote whichever number they prefer.
REPLACE_HOLDOUT = False

# %%
study = open_study(CASE_STUDY_ID, execution_tier=EXECUTION_TIER, workspace=WORKSPACE or None)
CASE_DIR = get_case_study_dir(CASE_STUDY_ID)


def _delete_holdout_generation(case_dir, prediction_hash):
    """Remove one holdout prediction set and everything registered against it.

    The rows go rather than being marked, because a holdout evaluation that is still
    readable is still a number someone can quote, and the point of removing it is that it
    should not be one.

    The cascade lives in `case_studies/utils/registry/maintenance.py` and derives the child
    tables from the schema. The version this replaces listed them by hand and was already
    missing `cohort_metrics.leader_hash`, which with foreign keys enabled aborts the delete
    rather than orphaning a row - so on any registry with a cohort this function raised
    instead of deleting, and the generation it was called on stayed.
    """
    deleted = delete_prediction_generation(case_dir / "run_log" / "registry.db", prediction_hash)
    for table, n in sorted(deleted.items()):
        print(f"  deleted {n:>3} from {table}")


# %% [markdown]
# ## 1. Which configuration the holdout runs
#
# The holdout runs the configuration the case study reports, resolved through the same
# `resolve_solvent_carrier` [`14_costs`](14_costs.ipynb) uses. Resolving it again here
# rather than passing it along is deliberate: the two notebooks must agree by construction,
# and a hash written down in one and read in the other agrees only until the sweep is
# rebuilt.
#
# Nothing about the holdout enters this choice. The selected configuration is the validation
# rank-1, and it was fixed before this notebook ran.

# %%
carrier = resolve_solvent_carrier(CASE_STUDY_ID)
print(
    f"Selected configuration: {carrier['val_backtest_hash']}  stage={carrier['val_stage']}  "
    f"family={carrier['family']}  config={carrier['config_name']}  "
    f"label={carrier['label']}"
)
print(
    f"  validation Sharpe {carrier['val_sharpe']:.3f}, max drawdown {carrier['max_drawdown']:.3f}"
)
print(f"  fitted by training run {carrier['training_hash']}")

# %% [markdown]
# The checkpoint is part of the configuration. This family publishes a prediction set per boosting
# iteration on a declared schedule, and the selected configuration's prediction set names one of
# them - so refitting without it would produce a model at the end of training rather than the one
# that was ranked.

# %%
validation_prediction = study.results.open(carrier["val_prediction_hash"])
prediction_record = validation_prediction.registry_record()
CHECKPOINT_KIND = prediction_record["checkpoint_kind"]
CHECKPOINT_VALUE = prediction_record["checkpoint_value"]
print(f"Checkpoint: {CHECKPOINT_KIND}={CHECKPOINT_VALUE}")

# %% [markdown]
# ## 2. The window, and the model that is allowed to see it
#
# The holdout window is not a choice made here. It is `evaluation.holdout_start` and
# `evaluation.holdout_end` from the case study's own `setup.yaml`, read through the same
# `canonical_window` the fold derivation and the backtest slice both go through, so the
# three cannot disagree.
#
# The training interval is everything available before that window, bounded above by a
# label buffer. The buffer is what stops the last training label's outcome from resolving
# inside the holdout: this case study dates each row by the month the return was earned,
# so a monthly label observed at the end of December is already realised, and the buffer
# is one observation rather than a horizon's worth. A zero gap would be a leak, not a
# conservative choice, so the derivation refuses to default it.
#
# Everything else about the configuration is carried across unchanged, and the fields that
# cannot be - the eligibility manifest, and any parameter this family resolves from a
# fold's own training rows - are recomputed against the holdout fold. Carrying those
# forward would fit a model keyed to the validation folds and call it a retrain.

# %%
observation_timeline = (
    pl.read_parquet(study.root / "labels" / f"{carrier['label']}.parquet")
    .get_column("timestamp")
    .unique()
    .sort()
    .to_list()
)
validation_spec = study.results.open(carrier["training_hash"]).spec()
holdout_spec = build_holdout_training_spec(
    study,
    validation_spec,
    timeline=observation_timeline,
    case_study=CASE_STUDY_ID,
)

fold = holdout_spec["computation"]["cv"]["folds"][0]
print(f"Holdout fold {fold['fold']}")
print(f"  trains  {fold['train_start']} -> {fold['train_end']}")
print(f"  predicts {fold['val_start']} -> {fold['val_end']}")
print(f"  label buffer: {holdout_spec['computation']['cv']['request']['label_buffer']}")

# The validation folds are what the buffer is measured against, and the last of them ends
# before the holdout opens. Printing both is what lets a reader check the gap rather than
# take it on the derivation's word.
validation_folds = validation_spec["computation"]["cv"]["folds"]
latest_validation_end = max(str(entry["val_end"]) for entry in validation_folds)
print(f"Validation folds: {len(validation_folds)}, latest evaluation end {latest_validation_end}")
print(f"Holdout training ends {fold['train_end']}, holdout opens {fold['val_start']}")

# %% [markdown]
# ## 3. Fit, and register the predictions
#
# `reconstruct_locked_model_request` builds the request from the spec above. Its name
# comes from the locked holdout path this case study no longer uses; it takes a training
# specification and a checkpoint, not a lock, and it is used here because it is the one
# call that refuses a request that is not exactly the spec it was handed - the training
# identity, the checkpoint schedule, the feature lineage and the runtime parameters are
# all checked before anything is fitted.
#
# The training identity below is new. It has to be: it covers the CV interval, and the
# holdout fold is not one of the validation folds. A run that came back with the
# validation training hash would mean the refit did not happen.
#
# **The window carries one configuration at a time.** The holdout is re-runnable, and that
# is not the same as free: every configuration evaluated on it is another look at a period
# the case study reports as unseen, and two evaluated quietly would make that report false.
#
# So the check below is on the selected configuration rather than on the notebook, and it has
# exactly two outcomes. With the selected configuration unchanged this is an idempotent replay: the
# derivation is deterministic and the training identity covers it, so the same identity comes back
# and the fit is served from the registry. With the selected configuration changed it refuses,
# names both configurations, and stops.
#
# `REPLACE_HOLDOUT` is the only way past that, and it is a replacement rather than an
# addition: the superseded generation's rows are deleted, so the registry never holds two
# refits of the holdout window and no downstream resolver has to choose between them.
# Deleting is what makes the earlier evaluation cost something to discard. It is also the
# only honest shape - a run that had been observed and then quietly kept alongside its
# replacement would let a reader take whichever number they preferred.

# %%
holdout_training_hash = training_hash_from_spec(holdout_spec)
this_generation = (holdout_training_hash, (CHECKPOINT_KIND, CHECKPOINT_VALUE))
retire = holdout_generations_to_retire(CASE_DIR, this_generation=this_generation)
# A row whose training run records no CV split cannot be shown either way, and deleting on
# that would discard a result nothing has established is wrong. It stops the run instead.
if retire.unattributable:
    raise RuntimeError(
        "the holdout window carries prediction sets whose training runs record no CV split, "
        "so whether they were refitted for the holdout cannot be established: "
        + ", ".join(
            f"{row['prediction_hash']} (training {row['training_hash']})"
            for row in retire.unattributable
        )
        + ". Establish what produced them before registering another evaluation on the same "
        "window; this notebook will not delete a row it cannot show is not a holdout result."
    )
# A row whose training run declares a non-holdout CV may not be reported as a holdout
# result, and it is also not something to delete unattended: `generate_holdout` refits on a
# holdout fold and then registers the predictions under the VALIDATION training identity, so
# this record covers both a validation-fitted model published over the window and a real
# refit filed under the wrong identity. Nothing owned this before - the filter here was
# `row["refitted"]`, which made exactly these rows invisible to the refusal and to
# everything after it.
if retire.not_out_of_sample and not REPLACE_HOLDOUT:
    raise RuntimeError(
        "the holdout window carries prediction sets whose training runs declare a CV split "
        "other than the holdout: "
        + ", ".join(
            f"{row['prediction_hash']} ({row['config_name']}, training {row['training_hash']})"
            for row in retire.not_out_of_sample
        )
        + ". Each is either a validation-fitted model published over the window, which is "
        "not an out-of-sample result, or a refit registered under its validation training "
        "identity, which `20_strategy_synthesis/holdout.py::generate_holdout` produces - and "
        "the registry cannot tell those apart. Establish which, then set REPLACE_HOLDOUT="
        "True to remove it, or leave it and resolve the identity instead."
    )
for row in retire.not_out_of_sample:
    print(
        f"REMOVING {row['prediction_hash']} ({row['config_name']}, training "
        f"{row['training_hash']}): its training run declares a non-holdout CV, so it is not "
        "reportable as a holdout evaluation under the identity it carries"
    )
    _delete_holdout_generation(CASE_DIR, row["prediction_hash"])
superseded = list(retire.superseded)
if superseded and not REPLACE_HOLDOUT:
    raise RuntimeError(
        "the holdout window already carries a refit of a different configuration: "
        + ", ".join(
            f"{row['prediction_hash']} ({row['config_name']}, training {row['training_hash']})"
            for row in superseded
        )
        + f". This run would evaluate {carrier['config_name']} (training "
        f"{holdout_training_hash}, checkpoint {CHECKPOINT_KIND}={CHECKPOINT_VALUE}) on the "
        "same window. Set REPLACE_HOLDOUT=True to discard the earlier generation, or leave "
        "the selection where it was."
    )
for row in superseded:
    print(f"REPLACING holdout generation {row['prediction_hash']} ({row['config_name']})")
    _delete_holdout_generation(CASE_DIR, row["prediction_hash"])

# %% tags=["results"]
request = reconstruct_locked_model_request(
    study,
    holdout_spec,
    checkpoint_kind=CHECKPOINT_KIND,
    checkpoint_value=CHECKPOINT_VALUE,
)
model_run = request.run()
holdout_prediction = model_run.predictions[0]

if model_run.training.hash == carrier["training_hash"]:
    raise RuntimeError(
        "the holdout refit produced the validation training identity "
        f"{carrier['training_hash']}, which means it did not refit"
    )
print(f"Holdout training run:  {model_run.training.hash}")
print(f"Holdout prediction set: {holdout_prediction.hash}")

# %% [markdown]
# What the prediction set covers, read back from the registry rather than from the
# request. The two agree only if the fit published what it declared, and the count is the
# one number a reader can check the window against: a monthly panel over one year is
# twelve decision dates, and the number of rows is those dates times the names eligible on
# each.

# %% tags=["results"]
record = holdout_prediction.registry_record()
predictions = holdout_prediction.load()
print(
    f"split={record['split']}  checkpoint={record['checkpoint_kind']}={record['checkpoint_value']}"
)
print(f"rows={predictions.height:,}  dates={predictions['timestamp'].n_unique()}")
print(
    f"  {predictions['timestamp'].min()} -> {predictions['timestamp'].max()}, "
    f"{predictions['symbol'].n_unique():,} names"
)

# %% [markdown]
# The registry now holds more than one holdout prediction set for this case study, and
# only one of them was fitted on data that ends before the window. The other is the
# defective generation this notebook replaces: it carries the validation training identity,
# which is how it was found. Both are listed rather than one silently preferred, because
# the registry is immutable and a reader looking at it later will see both.

# %% tags=["results"]
for row in registered_holdout_generations(CASE_DIR):
    note = (
        "refitted for the holdout" if row["refitted"] else "VALIDATION-FITTED - not out of sample"
    )
    print(
        f"  {row['prediction_hash']}  training={row['training_hash']}  {row['config_name']}  {note}"
    )

# %% [markdown]
# ## What this notebook establishes, and what it does not
#
# It establishes one thing: a prediction set over the holdout window, produced by the
# configuration this case study selected, fitted on data that ends before the window
# opens. That is a precondition for an out-of-sample claim, not the claim itself. Nothing
# here says whether the predictions are any good - they have not been scored, sized or
# traded.
#
# It does not make the holdout a fresh test in the strict sense. The configuration reached
# this notebook through a selection made on the validation folds, and this window is being
# used once per configuration that gets here. What it does remove is the specific
# circularity of scoring a validation-fitted model on the period meant to judge it.
#
# The holdout is re-runnable. If a later pass finds the selection was wrong, the answer is
# to delete this generation and produce another, not to treat the first as spent.
#
# **Next:** [`16_holdout_backtest`](16_holdout_backtest.ipynb).

```

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: MIT

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.