مواد پر جائیں
لائبریری کی تمام دستاویزات

لیکج کنٹرول والے ہولڈ آؤٹ کے لیے ETF ماڈلز کی ازسرِنو تربیت

نوٹ بک Machine Learning for Trading

خلاصہ

یہ کیس اسٹڈی بتاتی ہے کہ توثیقی ڈیٹا پر ماڈل اور حکمتِ عملی کے انتخاب کے بعد ہولڈ آؤٹ مدت کے لیے پیش گوئیاں کیسے تیار کی جائیں۔ یہ پہلے سے منتخب ترتیب طے کرتی، اس کی تربیتی تخصیص دوبارہ بناتی اور ہولڈ آؤٹ شروع ہونے سے پہلے دستیاب ڈیٹا ہی سے اسے دوبارہ فٹ کرتی ہے۔ چونکہ لیبل مستقبل کے منافع ظاہر کرتے ہیں، اس لیے تربیتی وقفہ تشخیصی مدت سے پورے لیبل افق جتنا پہلے ختم ہوتا ہے تاکہ اس مدت کے نتائج ماڈل فٹنگ میں شامل نہ ہوں۔

طریقۂ کار جانچتا ہے کہ ہولڈ آؤٹ فٹ کی تربیتی شناخت الگ ہے، اس کی پیش گوئیاں درج کرتا ہے اور بتاتا ہے کہ درج شدہ پیش گوئیوں کے مجموعے واقعی دوبارہ فٹ کیے گئے یا نہیں۔ یہ منتخب ترتیب کو طے شدہ سمجھتا اور اسی مدت میں دوسری ترتیب سے انکار کرتا ہے، کیونکہ ہولڈ آؤٹ نتائج دیکھنے کے بعد انتخاب بدلنا جانچ کو متاثر کرے گا۔ یہ نوٹ بک پیش گوئیوں کا ماخذ درج کرتی ہے، حکمتِ عملی کی کارکردگی نہیں: اسکورنگ، حجم بندی اور ٹریڈنگ بعد کے مراحل پر ہیں۔ ہولڈ آؤٹ بھی توثیق پر مبنی انتخاب کے بعد آتا ہے، اس لیے وہ پہلے کے انتخابی عمل کو ختم نہیں کرتا۔

اہم خیالات

  • ہولڈ آؤٹ ماڈل کو ہولڈ آؤٹ سے پہلے کے ڈیٹا پر دوبارہ فٹ کرنا چاہیے، توثیق پر فٹ ماڈل دوبارہ استعمال نہیں کرنا چاہیے۔
  • مستقبل کے منافع والے لیبلز کے لیے وقفہ رکھیں تاکہ تربیتی نتائج تشخیصی مدت تک نہ پھیلیں۔
  • تربیتی شناختیں اور رجسٹری ریکارڈ بتاتے ہیں کہ کن ڈیٹا اور ترتیب سے پیش گوئیاں تیار ہوئیں۔
  • صرف پیش گوئیاں ان کے معیار یا ٹریڈ کے قابل حکمتِ عملی کو ثابت نہیں کرتیں؛ بیک ٹیسٹنگ اور تجزیہ پھر بھی ضروری ہیں۔
  • نئی منتخب ترتیب کی جانچ کے لیے ہولڈ آؤٹ دوبارہ استعمال کرنے سے اس کی آؤٹ آف سیمپل قدر کمزور ہوتی ہے۔

ٹیگز

مکمل متن
# ETFs: Holdout Predictions


# ETFs: Holdout Predictions

**Chapter 20 - Out-of-sample evaluation**

Every number in this case study so far was measured on the validation folds, and every
choice was made by looking at them: which model family, which configuration, how many
funds to hold, how to size them, which risk control to overlay, what to charge. A result
selected that way cannot also be evidence that the selection was sound - the ranking and
the evidence would be the same measurement.

The holdout is the window nothing has been selected on. This notebook fits the selected
configuration on the history available before that window opens and writes its
predictions over it. [`19_holdout_backtest`](19_holdout_backtest.ipynb) turns those
predictions into a return series with the sizing and the overlay the case study settled
on, and [`20_strategy_analysis`](20_strategy_analysis.ipynb) reads both back.

**What this notebook is careful about**

A holdout prediction is not the validation model scored on a later window. Section 2
fits again, over a training interval that ends before the window opens, and the new
training identity is what makes the refit visible rather than asserted: the identity
covers the CV interval, so a run that came back with the validation training hash would
mean no refit happened. The check is in section 3 and it raises.

**Prerequisites:** [`17_costs`](17_costs.ipynb), which is the last stage that selects.

**Scope:** one training run and one prediction set. No backtest, no selection, no
comparison - those are 19 and 20.

```python
"""ETFs: Holdout Predictions."""

import polars as pl

from case_studies.research import open_study
from case_studies.research.holdout import build_holdout_training_spec
from case_studies.research.models import reconstruct_locked_model_request
from case_studies.utils.registry import training_hash_from_spec
from case_studies.utils.strategy_analysis import (
    holdout_generations_to_retire,
    registered_holdout_generations,
    resolve_solvent_carrier,
)
from utils.paths import get_case_study_dir
```

```python
CASE_STUDY_ID = "etfs"
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""
```

```python
study = open_study(CASE_STUDY_ID, execution_tier=EXECUTION_TIER, workspace=WORKSPACE or None)
CASE_DIR = get_case_study_dir(CASE_STUDY_ID)
```

## 1. Which configuration the holdout runs

The holdout runs the configuration the case study reports, resolved through the same
`resolve_solvent_carrier` [`17_costs`](17_costs.ipynb) prices. Resolving it again here
rather than passing it along is deliberate: the two notebooks must agree by construction,
and a hash written down in one and read in the other agrees only until the sweep is
rebuilt.

Nothing about the holdout enters this choice. The selected configuration is the cross-stage
validation rank-1 - which for this case study is a risk-overlay run, not the allocation leader -
and it was fixed before this notebook ran.

```python
carrier = resolve_solvent_carrier(CASE_STUDY_ID)
print(
    f"Selected configuration: {carrier['val_backtest_hash']}  stage={carrier['val_stage']}  "
    f"family={carrier['family']}  config={carrier['config_name']}  "
    f"label={carrier['label']}"
)
print(
    f"  validation Sharpe {carrier['val_sharpe']:.3f}, max drawdown {carrier['max_drawdown']:.3f}"
)
print(f"  fitted by training run {carrier['training_hash']}")
```

The checkpoint is part of the configuration. Where a family publishes a prediction set per
checkpoint on a declared schedule, the selected configuration's prediction set names one of them,
and refitting without it would produce a model at the end of training rather than the one that
was ranked. A family with no checkpoint dimension stores NULL in both columns and carries that
NULL through unchanged.

```python
validation_prediction = study.results.open(carrier["val_prediction_hash"])
prediction_record = validation_prediction.registry_record()
CHECKPOINT_KIND = prediction_record["checkpoint_kind"]
CHECKPOINT_VALUE = prediction_record["checkpoint_value"]
print(f"Checkpoint: {CHECKPOINT_KIND}={CHECKPOINT_VALUE}")
```

## 2. The window, and the model that is allowed to see it

The holdout window is not a choice made here. It is `evaluation.holdout_start` and
`evaluation.holdout_end` from this case study's own `setup.yaml`, read through the same
`canonical_window` the fold derivation and the backtest slice both go through, so the three
cannot disagree.

The training interval is everything available before that window, bounded above by a label
buffer. **For this case study the buffer is doing real work.** The label is a 21-day forward
return, so a row dated `t` records an outcome that is not known until `t + 21` sessions - and
a training set running to the day the window opens would be fitted on labels whose returns
resolve inside it. The buffer is therefore a horizon's worth rather than a formality, and the
derivation refuses to default it: a zero gap here would be a leak.

Everything else about the configuration is carried across unchanged, and the fields that
cannot be - the eligibility manifest, and any parameter this family resolves from a fold's
own training rows - are recomputed against the holdout fold. Carrying those forward would fit
a model keyed to the validation folds and call it a retrain.

```python
observation_timeline = (
    pl.read_parquet(study.root / "labels" / f"{carrier['label']}.parquet")
    .get_column("timestamp")
    .unique()
    .sort()
    .to_list()
)
validation_spec = study.results.open(carrier["training_hash"]).spec()
holdout_spec = build_holdout_training_spec(
    study,
    validation_spec,
    timeline=observation_timeline,
    case_study=CASE_STUDY_ID,
)

fold = holdout_spec["computation"]["cv"]["folds"][0]
print(f"Holdout fold {fold['fold']}")
print(f"  trains  {fold['train_start']} -> {fold['train_end']}")
print(f"  predicts {fold['val_start']} -> {fold['val_end']}")
print(f"  label buffer: {holdout_spec['computation']['cv']['request']['label_buffer']}")

# The validation folds are what the buffer is measured against, and the last of them ends
# before the holdout opens. Printing both is what lets a reader check the gap rather than take
# it on the derivation's word.
validation_folds = validation_spec["computation"]["cv"]["folds"]
latest_validation_end = max(str(entry["val_end"]) for entry in validation_folds)
print(f"Validation folds: {len(validation_folds)}, latest evaluation end {latest_validation_end}")
print(f"Holdout training ends {fold['train_end']}, holdout opens {fold['val_start']}")
```

## 3. Fit, and register the predictions

`reconstruct_locked_model_request` builds the request from the spec above. Its name comes
from a locked holdout path this case study does not use; it takes a training specification
and a checkpoint, not a lock, and it is used here because it is the one call that refuses a
request that is not exactly the spec it was handed - the training identity, the checkpoint
schedule, the feature lineage and the runtime parameters are all checked before anything is
fitted.

The training identity below is new. It has to be: it covers the CV interval, and the holdout
fold is not one of the validation folds. A run that came back with the validation training
hash would mean the refit did not happen, so that is checked rather than assumed.

**The window carries one configuration, and this notebook has no way past that.** The check below
is on the selected configuration rather than on the notebook, and it has exactly two outcomes.
With the selected configuration unchanged this is an idempotent replay: the derivation is
deterministic and the training identity covers it, so the same identity comes back and the fit is
served from the registry, which is why re-running the notebook is free and safe. With the
selected configuration changed it refuses, names both configurations, and stops.

It refuses rather than offering a replacement switch, and the reason is that a replacement
would not be one. Deleting the earlier generation's rows does not undo having observed its
result: the selection that produced the new configuration may have been informed by the old
holdout number, and no deletion reaches that. A switch here would let the case study take a
second look at the window while leaving a registry that shows only one, which is the
specific thing that would make the out-of-sample claim false rather than merely weak.

```python
holdout_training_hash = training_hash_from_spec(holdout_spec)
this_generation = (holdout_training_hash, (CHECKPOINT_KIND, CHECKPOINT_VALUE))
retire = holdout_generations_to_retire(CASE_DIR, this_generation=this_generation)
# A row whose training run records no CV split cannot be shown either way, and deleting on
# that would discard a result nothing has established is wrong. It stops the run instead.
if retire.unattributable:
    raise RuntimeError(
        "the holdout window carries prediction sets whose training runs record no CV split, "
        "so whether they were refitted for the holdout cannot be established: "
        + ", ".join(
            f"{row['prediction_hash']} (training {row['training_hash']})"
            for row in retire.unattributable
        )
        + ". Establish what produced them before registering another evaluation on the same "
        "window; this notebook will not delete a row it cannot show is not a holdout result."
    )
# A row whose training run declares a non-holdout CV may not be reported as a holdout
# result, and it is also not something to delete unattended: `generate_holdout` refits on a
# holdout fold and then registers the predictions under the VALIDATION training identity, so
# this record covers both a validation-fitted model published over the window and a real
# refit filed under the wrong identity. Nothing owned this before - the filter here was
# `row["refitted"]`, which made exactly these rows invisible to the refusal and to
# everything after it.
if retire.not_out_of_sample:
    raise RuntimeError(
        "the holdout window carries prediction sets whose training runs declare a CV split "
        "other than the holdout: "
        + ", ".join(
            f"{row['prediction_hash']} ({row['config_name']}, training {row['training_hash']})"
            for row in retire.not_out_of_sample
        )
        + ". Each is either a validation-fitted model published over the window, which is "
        "not an out-of-sample result, or a refit registered under its validation training "
        "identity, which the retired `20_strategy_synthesis/holdout.py::generate_holdout` "
        "wrote until it was deleted on 2026-09-12 - and "
        "the registry cannot tell those apart. This notebook has no way past that: "
        "establish which it is and resolve it through the registry's own lifecycle, which "
        "records that the row was retired."
    )
superseded = list(retire.superseded)
if superseded:
    raise RuntimeError(
        "the holdout window already carries a refit of a different configuration: "
        + ", ".join(
            f"{row['prediction_hash']} ({row['config_name']}, training {row['training_hash']})"
            for row in superseded
        )
        + f". This run would evaluate {carrier['config_name']} (training "
        f"{holdout_training_hash}, checkpoint {CHECKPOINT_KIND}={CHECKPOINT_VALUE}) on the "
        "same window, which would be a second configuration measured on a period this case "
        "study reports as unseen. This notebook has no way past that: deleting the earlier "
        "generation would not undo having observed it, and the selection bias it introduces "
        "is not removed by removing the rows. Either leave the selection where it was, or "
        "retire the earlier evaluation through the registry's own lifecycle, which records "
        "that a second look was taken."
    )
```

```python
request = reconstruct_locked_model_request(
    study,
    holdout_spec,
    checkpoint_kind=CHECKPOINT_KIND,
    checkpoint_value=CHECKPOINT_VALUE,
)
model_run = request.run()
holdout_prediction = model_run.predictions[0]

if model_run.training.hash == carrier["training_hash"]:
    raise RuntimeError(
        "the holdout refit produced the validation training identity "
        f"{carrier['training_hash']}, which means it did not refit"
    )
print(f"Holdout training run:  {model_run.training.hash}")
print(f"Holdout prediction set: {holdout_prediction.hash}")
```

What the prediction set covers, read back from the registry rather than from the request.
The two agree only if the fit published what it declared, and the counts are what a reader
checks the window against: this is a daily panel, so the date count is trading sessions in
the window and the row count is those sessions times the funds eligible on each.

```python
record = holdout_prediction.registry_record()
predictions = holdout_prediction.load()
print(
    f"split={record['split']}  checkpoint={record['checkpoint_kind']}={record['checkpoint_value']}"
)
print(f"rows={predictions.height:,}  sessions={predictions['timestamp'].n_unique():,}")
print(
    f"  {predictions['timestamp'].min()} -> {predictions['timestamp'].max()}, "
    f"{predictions['symbol'].n_unique():,} funds"
)
```

Every holdout prediction set the registry holds, and whether the model behind it was
refitted for the window. All of them are listed rather than one silently preferred, because
the registry is immutable and a reader looking at it later will see whatever is there. A row
marked VALIDATION-FITTED is not an out-of-sample result whatever its numbers say.

```python
for row in registered_holdout_generations(CASE_DIR):
    note = (
        "refitted for the holdout" if row["refitted"] else "VALIDATION-FITTED - not out of sample"
    )
    print(
        f"  {row['prediction_hash']}  training={row['training_hash']}  {row['config_name']}  {note}"
    )
```

## What this notebook establishes, and what it does not

It establishes one thing: a prediction set over the holdout window, produced by the
configuration this case study selected, fitted on data that ends a full label horizon before
the window opens. That is a precondition for an out-of-sample claim, not the claim itself.
Nothing here says whether the predictions are any good - they have not been scored, sized or
traded.

It does not make the holdout a fresh test in the strict sense. The configuration reached this
notebook through a selection made on the validation folds, and this window is being used once
per configuration that gets here. What it does remove is the specific circularity of scoring
a validation-fitted model on the period meant to judge it.

Re-running this notebook is free: the same configuration re-derives the same training identity and
the fit is served from the registry. Evaluating a DIFFERENT configuration is not, and is
refused here. If a later pass finds the selection was wrong, that is a question for the
registry's lifecycle, which records that a second look was taken - not something to settle by
deleting rows until the registry agrees.

**Next:** [`19_holdout_backtest`](19_holdout_backtest.ipynb).

ماخذ کا حوالہ دیتے ہوئے مکمل متن دکھایا گیا ہے، ماخذ کے لائسنس کے تحت۔ لائسنس: MIT

یہ خلاصہ اصل ماخذ سے Stratmill کے تحقیقی ایجنٹ نے لکھا ہے؛ یہ ماخذ کی نقل نہیں۔