Holdout-Backtest einer festen Strategie interpretieren
Zusammenfassung
Dieses Notebook wendet eine zuvor ausgewählte NASDAQ-100-Handelskonfiguration auf Prognosen für ein unangetastetes Holdout-Zeitfenster an. Strategie, Portfoliokonzentration, Allokationsmethode, Rebalancing-Zeitplan, Risiko-Overlay und Annahme zu Handelskosten stehen vorab fest; das Notebook trifft keine neuen Entscheidungen. Es registriert den resultierenden Backtest und berichtet Rendite- und Handelsstatistiken neben den Validierungswerten.
Die zentrale Erkenntnis lautet: Ein Out-of-Sample-Ergebnis ist nur dann aussagekräftig, wenn die Entscheidungen unverändert bleiben, nachdem das Holdout betrachtet wurde. Die Validierungsleistung ist das beste Ergebnis aus einer Suche über viele Backtests, während die Holdout-Leistung eine einzelne Messung über einen kürzeren Zeitraum ist. Ihre Differenz allein kann daher weder einen Strategieabbau noch einen Selektionsbias belegen. Für diesen Vergleich verweist das Notebook auf eine separate Analyse.
Das Holdout belegt die Leistung über einen Zeitraum, aber nicht, dass die ausgewählte Konfiguration die beste war oder ihre Ergebnisse anhalten werden. Das kurze Zeitfenster lässt erhebliche Unsicherheit bestehen, und das Validierungsergebnis kann optimistisch ausfallen, weil es aus einer großen Kandidatenmenge ausgewählt wurde.
Kernaussagen
- Ein Holdout-Backtest sollte eine vollständig festgelegte Strategiekonfiguration ohne weitere Auswahl anwenden.
- Die aus vielen Kandidaten ausgewählte Validierungsleistung ist nicht direkt mit einer einzelnen, kürzeren Holdout-Messung vergleichbar.
- Eine Leistungslücke allein kann einen Strategieabbau nicht von gewöhnlichen Schwankungen in einer kurzen Stichprobe unterscheiden.
- Das Holdout misst eine Konfiguration über einen nicht beobachteten Zeitraum; es beweist nicht, dass diese Konfiguration die beste Wahl war.
Schlagwörter
Volltext
# NASDAQ-100 Microstructure: Holdout Backtest
# NASDAQ-100 Microstructure: Holdout Backtest
**Chapter 20 - Out-of-sample evaluation**
[`18_holdout_predictions`](18_holdout_predictions.ipynb) refitted the selected
configuration on the history before the holdout window and wrote its predictions over it.
This notebook trades them, with the concentration, the allocator, the risk overlay and the
cost assumption the rest of the case study used, and registers the result.
Nothing is chosen here. The predictions, the allocator, the concentration, the rebalance
cadence, the overlay and the charge all arrive fixed from earlier notebooks, and the only
thing this notebook decides is that they are applied unchanged. That is the whole design: a
holdout result is worth something exactly to the extent that no decision was made after
seeing it, and every knob left open here would be a decision.
The comparison to validation is printed but not interpreted. The validation figure is the
maximum of a ranking over more than a thousand backtests and the holdout figure is one
measurement over a much shorter window; what can be said about the gap between them is
[`20_strategy_analysis`](20_strategy_analysis.ipynb)'s subject, with the intervals to say
it.
**Prerequisites:** [`18_holdout_predictions`](18_holdout_predictions.ipynb).
**Scope:** one backtest. No selection, no comparison beyond a printed pair.
```python
"""NASDAQ-100 Microstructure: Holdout Backtest."""
import json
import sqlite3
import polars as pl
from case_studies.research import open_study
from case_studies.research.holdout import build_holdout_training_spec
from case_studies.utils.backtest_loaders import get_backtest_config, load_backtest_prices_for
from case_studies.utils.backtest_presets import ensure_backtest_spec, strategy_view
from case_studies.utils.backtest_runner import resolved_allow_short_selling, run_backtest
from case_studies.utils.conformal import (
compute_holdout_conformal_widths,
ensure_conformal_calibration_identity,
holdout_conformal_embargo_steps,
)
from case_studies.utils.registry import (
backtest_run_status,
read_predictions,
training_hash_from_spec,
)
from case_studies.utils.strategy_analysis import resolve_solvent_carrier
from utils.paths import get_case_study_dir
```
```python
CASE_STUDY_ID = "nasdaq100_microstructure"
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""
MAX_SYMBOLS = 0
```
```python
study = open_study(CASE_STUDY_ID, execution_tier=EXECUTION_TIER, workspace=WORKSPACE or None)
CASE_DIR = get_case_study_dir(CASE_STUDY_ID)
bt_config = get_backtest_config(CASE_STUDY_ID)
# `MAX_SYMBOLS` reduces the price panel this run trades and reaches `backtest_hash` through
# nothing, so a reduced run and a full run over the same holdout predictions hash alike and
# the second is served the first's result (ml4t/agent-workspace#911). `14_backtest` and
# `15_portfolio_management` give a reduced run an identity of its own by declaring the traded
# universe into the spec they BUILD; this notebook carries the selected configuration's spec
# forward through `ensure_backtest_spec`, which has no such parameter, so there is no identity
# to give one here. That leaves refusal as the only correct answer on the canonical tier, and
# it matters more here than anywhere else in the case study: this is the one window the study
# reports as unseen, it carries one backtest, and a narrowed result registered against it
# could not be distinguished afterwards from the declared portfolio's.
if EXECUTION_TIER == "canonical" and MAX_SYMBOLS:
raise ValueError(
"MAX_SYMBOLS narrows the universe this run trades, which makes it a different "
"portfolio from the declared one and gives it its own backtest identity "
"(ml4t/agent-workspace#911). A canonical run trades the declared universe: set "
"MAX_SYMBOLS=0, or run under EXECUTION_TIER='preview' with a WORKSPACE."
)
def _registered_holdout_backtests(case_dir, prediction_hash):
"""The backtest hashes already registered against one holdout prediction set."""
with sqlite3.connect(str(case_dir / "run_log" / "registry.db")) as conn:
rows = conn.execute(
"SELECT backtest_hash FROM backtest_runs WHERE prediction_hash = ? "
"ORDER BY backtest_hash",
(prediction_hash,),
).fetchall()
return [backtest_hash for (backtest_hash,) in rows]
```
## 1. The configuration, and the predictions it produced on the holdout
The selected configuration is resolved the same way [`17_costs`](17_costs.ipynb) and
[`18_holdout_predictions`](18_holdout_predictions.ipynb) resolve it, so all three run the
same configuration by construction rather than by a hash copied between them.
Which holdout prediction set belongs to it is derived rather than searched for. Re-deriving the
holdout training specification reproduces the training identity 18 registered - the derivation is
deterministic and the identity covers it - so the prediction set is looked up by that identity
and the selected configuration's checkpoint. A search over holdout prediction sets would have to
guess which one belonged to this configuration.
```python
carrier = resolve_solvent_carrier(CASE_STUDY_ID)
LABEL = carrier["label"]
validation_prediction_record = study.results.open(carrier["val_prediction_hash"]).registry_record()
holdout_spec = build_holdout_training_spec(
study,
study.results.open(carrier["training_hash"]).spec(),
timeline=(
pl.read_parquet(study.root / "labels" / f"{LABEL}.parquet")
.get_column("timestamp")
.unique()
.sort()
.to_list()
),
case_study=CASE_STUDY_ID,
)
holdout_training_hash = training_hash_from_spec(holdout_spec)
with sqlite3.connect(str(CASE_DIR / "run_log" / "registry.db")) as conn:
match = conn.execute(
"""
SELECT prediction_hash FROM prediction_sets
WHERE split = 'holdout' AND training_hash = ?
AND checkpoint_kind IS ? AND checkpoint_value IS ?
""",
(
holdout_training_hash,
validation_prediction_record["checkpoint_kind"],
validation_prediction_record["checkpoint_value"],
),
).fetchone()
if match is None:
# Two causes, and they call for opposite actions, so the message has to tell them apart.
# 18 not having run is the obvious one. The other is that the carrier MOVED between 18 and
# this notebook: both resolve independently and by design, and this case study declares no
# backtest population, so `published_members_at(member_kind="backtest")` is None and no
# backtest row is ever retired from the pool. A sweep that registers a higher-Sharpe cell
# between the two runs therefore changes what `resolve_solvent_carrier` returns, this
# notebook re-derives a different training identity, and the lookup misses.
#
# Sending that case to "run 18 first" points at the wrong thing. 18 would refuse it - its
# `holdout_generations_to_retire` check exists for exactly this - but only after the reader
# has spent a cycle being told to do the thing that cannot work.
with sqlite3.connect(str(CASE_DIR / "run_log" / "registry.db")) as conn:
registered = conn.execute(
"""
SELECT t.config_name, t.label, p.training_hash,
p.checkpoint_kind, p.checkpoint_value
FROM prediction_sets p JOIN training_runs t ON t.training_hash = p.training_hash
WHERE p.split = 'holdout'
ORDER BY p.created_at
"""
).fetchall()
# Checkpoint is part of the generation, not a detail below it: 18 keys its retirement
# check on `(training_hash, (checkpoint_kind, checkpoint_value))`, and the lookup above
# matches on all three. So a set for this configuration at a DIFFERENT checkpoint is a
# different configuration and belongs in the moved-selection branch, not the replay one.
_ckpt = (
validation_prediction_record["checkpoint_kind"],
validation_prediction_record["checkpoint_value"],
)
same_config = [
(cfg, lab, th)
for cfg, lab, th, ck_kind, ck_value in registered
if cfg == carrier["config_name"] and lab == LABEL and (ck_kind, ck_value) == _ckpt
]
named = ", ".join(
f"{cfg} on {lab} at {ck_kind}={ck_value} ({th})"
for cfg, lab, th, ck_kind, ck_value in registered[:4]
)
if not registered:
msg = (
f"No holdout prediction set for training {holdout_training_hash}, and this registry "
"holds none at all. Run 18_holdout_predictions first; this notebook does not fit."
)
elif same_config:
# The configuration is the one that was fitted, under a training identity that is not
# the one this notebook just derived. Selection did not move; the DERIVATION did - the
# holdout spec, the label timeline it is built from, or something the identity hashes.
# That diagnosis differs from the branch below, which is why they are separated, but
# the action does not: 18 refuses in both cases. `holdout_generations_to_retire`
# skips only a row whose (training_hash, checkpoint) equals the generation about to
# be registered, so a row under the earlier identity is not skipped. It lands in one
# of the three retirement buckets by its training spec - `superseded` if it was a
# genuine refit, otherwise `not_out_of_sample` or `unattributable` - and each of the
# three raises in 18 before anything is fitted. The `superseded` message reads "a
# refit of a different configuration", which is 18's identity for a generation and
# not the config name.
msg = (
f"No holdout prediction set for training {holdout_training_hash}, but "
f"{carrier['family']}/{carrier['config_name']} on {LABEL} at "
f"{_ckpt[0]}={_ckpt[1]} - the configuration that resolves now - already has one under "
f"{', '.join(th for _, _, th in same_config)}. So the selection did not move and "
"its derived training identity did: the holdout spec, the label timeline it reads, "
"or an input the identity covers has changed since 18 ran. Re-running 18 on its "
"own will NOT recover this: it matches a registered generation on the exact "
"training hash and checkpoint, so the row under the earlier identity reads as a "
"superseded refit and 18 refuses before fitting. Either restore the derivation "
"the registered row was fitted under, or retire that generation through the "
"registry's own lifecycle - which records that the window was looked at twice - "
"and then re-run 18."
)
else:
msg = (
f"No holdout prediction set for training {holdout_training_hash}, but this registry "
f"holds {len(registered)} for other configurations: {named}. So "
"18_holdout_predictions has run and the selected configuration has MOVED since - it "
f"now resolves to {carrier['family']}/{carrier['config_name']} on {LABEL}. A "
"backtest row is never retired from the carrier pool, so a sweep registering a "
"higher-Sharpe cell between the two notebooks is enough to do it. Do NOT re-run 18 "
"to fit the new one: that spends a second holdout observation on a second "
"configuration, which is what its retirement check refuses. Decide which "
"configuration this case study carries."
)
raise RuntimeError(msg)
HOLDOUT_PREDICTION_HASH = match[0]
print(f"Selected configuration: {carrier['val_backtest_hash']} {carrier['config_name']} ({LABEL})")
print(f"Holdout training: {holdout_training_hash}")
print(f"Holdout prediction: {HOLDOUT_PREDICTION_HASH}")
```
## 2. Calibration, where the allocator needs it
An allocator that sizes by a conformal width is calibrated from errors the model has already
made, and on the holdout there are none to use: an error is usable only once the return it
measures has been realised, and every holdout return realises inside the window being evaluated.
Such a selected configuration takes its widths from the validation residuals of the validation
prediction set, which is what the allocator would have had standing at the start of the window.
The embargo matters for this case study's label and would not for every one. A residual
observed at `t` is not resolved until `t + h + 1min`, one bar PAST the label's horizon: the
entry leg is the VWAP of the bar after the decision, so a label at `t` consumes a quote at
`t + h + 1`. So the last residuals of the validation span reach into the holdout window, and
calibrating on them would size holdout positions with holdout price information.
The value comes from the reviewed table in `conformal.py`, which records the label BUFFER rather
than the horizon for exactly that reason - 6, 16 and 61 for `fwd_ret_5m`,
`fwd_ret_15m`/`fwd_dir_15m` and `fwd_ret_60m`. It is read from the selected configuration's label
rather than assumed from the primary, because the selected configuration is resolved across every
declared label. The unit is prediction-data steps, and on this panel a step is one MINUTE, not
one decision slot: the modal gap between adjacent prediction timestamps is 0:01:00 for all four
labels. Counting in decision slots is what an earlier version of that table did, and it embargoed
one minute where the label reaches sixteen.
The reach is minutes rather than the ETF study's three weeks, and it is the same defect at any
width: the residuals that cross the boundary are the ones nearest it, which are the ones a
calibration weights most.
The branch is here rather than assumed away because the selected configuration can change. It
does not fire for the one this case study currently reports, which sizes by inverse volatility
over a declared window and needs no calibration - so the line below prints that rather than
staying silent, which is what tells a reader the branch was evaluated.
The widths themselves are NOT written here. Writing them replaces the artifact an already
registered run was sized by, and the replacement guard in section 3 can still refuse this run
afterwards - which would leave the registered holdout pointing at a calibration that no
longer existed. The write is below the guard.
```python
allocation = strategy_view(json.loads(carrier["spec_json"])).get("allocation") or {}
NEEDS_CALIBRATION = allocation.get("method") == "conformal_weighted"
embargo_steps = holdout_conformal_embargo_steps(CASE_STUDY_ID, LABEL) if NEEDS_CALIBRATION else 0
if NEEDS_CALIBRATION:
print(f"Conformal configuration: embargo {embargo_steps} observation(s), widths written below.")
else:
print(f"Allocator {allocation.get('method', 'equal_weight')!r} needs no calibration.")
```
## 3. The backtest
The strategy specification is the selected configuration's own, re-pointed at the holdout
prediction set and the holdout price window. Nothing else about it changes - the concentration,
the allocator, the risk overlay and the commission and slippage levels are the ones `setup.yaml`
declares and every validation number in this case study was net of.
The run registers under `stage='holdout'`, which the registry derives from the prediction set's
split rather than from anything asserted here - and that derivation takes precedence over the
risk block the selected configuration carries, which would otherwise file this as another risk
overlay.
**The window carries one backtest**, for the same reason 18 lets it carry one prediction
generation. 18's guard is on the model - the training identity and the checkpoint - and it
cannot see this one: a changed allocator, overlay, cost level or calibration produces the
same holdout predictions and a different result from them. Like 18's, this guard refuses
rather than offering a replacement, because deleting a result that has been seen does not
unsee it.
The test is the backtest hash, not a field-by-field comparison. Every input that changes the
result is in that hash by construction, and a guard naming fields instead has to be right
about all of them. The hash is resolved before anything runs, so nothing is evaluated on the
holdout before the question is answered, and it comes from `backtest_run_status` - the call
the runner itself makes - because asking it is the only way to be sure the guard and the
runner agree about identity. The run asserts they still agree afterwards, because a guard
that had quietly stopped predicting the hash would let everything through while looking
correct.
```python
prices = load_backtest_prices_for(CASE_STUDY_ID, LABEL, split="holdout", max_symbols=MAX_SYMBOLS)
predictions = read_predictions(CASE_STUDY_ID, HOLDOUT_PREDICTION_HASH)
print(f"Prices: {len(prices):,} rows, {prices['symbol'].n_unique():,} symbols")
print(
f"Predictions: {predictions.height:,} rows, "
f"{predictions['timestamp'].n_unique():,} decision timestamps, "
f"{predictions['timestamp'].dt.date().n_unique():,} sessions"
)
spec = ensure_backtest_spec(
CASE_STUDY_ID,
bt_config,
json.loads(carrier["spec_json"]),
prices=prices,
prediction_hash=HOLDOUT_PREDICTION_HASH,
initial_cash=bt_config.initial_cash,
)
spec["chapter"] = "ch20"
# The embargo goes into the specification before anything hashes it. The widths are an input to
# this backtest and the embargo decides them, so two embargoes are two results and must not
# share an identity.
if NEEDS_CALIBRATION:
spec = ensure_conformal_calibration_identity(spec, holdout_embargo_steps=embargo_steps)
spec["backtest_config"]["account"]["allow_short_selling"] = resolved_allow_short_selling(spec, None)
prospective_hash = backtest_run_status(CASE_STUDY_ID, HOLDOUT_PREDICTION_HASH, spec).backtest_hash
superseded_backtests = sorted(
set(_registered_holdout_backtests(CASE_DIR, HOLDOUT_PREDICTION_HASH)) - {prospective_hash}
)
if superseded_backtests:
raise RuntimeError(
"the holdout window already carries a backtest of a different configuration: "
+ ", ".join(superseded_backtests)
+ f". This run would register {prospective_hash} and has not run. Same rule as "
"18_holdout_predictions: discarding the earlier result would not undo having observed "
"it, so there is no switch here. Leave the selection where it was, or retire the "
"earlier evaluation through the registry's lifecycle."
)
# The guard has passed, so this run will register and the widths it is sized by are the ones
# that belong beside this prediction set.
if NEEDS_CALIBRATION:
widths = compute_holdout_conformal_widths(
CASE_STUDY_ID,
carrier["val_prediction_hash"],
HOLDOUT_PREDICTION_HASH,
alpha=float(allocation.get("alpha", 0.2)),
min_calibration_n=int(allocation["min_calibration_n"]),
embargo_steps=embargo_steps,
write=True,
)
print(
f"Conformal widths: {widths.height:,} rows over "
f"{widths['symbol'].n_unique():,} names, embargo {embargo_steps} observation(s)"
)
result = run_backtest(
CASE_STUDY_ID,
HOLDOUT_PREDICTION_HASH,
spec,
prices=prices,
predictions=predictions,
label=LABEL,
register=True,
initial_cash=bt_config.initial_cash,
calendar=bt_config.calendar,
)
if result.backtest_hash != prospective_hash:
raise RuntimeError(
f"the guard predicted {prospective_hash} and the runner registered "
f"{result.backtest_hash}. The guard decides what may run on the holdout, so a guard "
"that no longer reproduces the runner's identity is not a smaller problem than the one "
"it was written for."
)
print(f"Holdout backtest: {result.backtest_hash}")
# The stage is checked rather than trusted. This configuration comes from the risk stage and its
# spec carries a risk block, and stage inference reads the prediction's split before that block - so
# a holdout run files as `holdout`. If that order ever changes, the whole out-of-sample result lands
# in `risk_overlay` and `20_strategy_analysis` finds no holdout at all, which is a failure four
# notebooks away from its cause.
with sqlite3.connect(str(CASE_DIR / "run_log" / "registry.db")) as conn:
registered_stage = conn.execute(
"SELECT stage FROM backtest_runs WHERE backtest_hash = ?", (result.backtest_hash,)
).fetchone()[0]
if registered_stage != "holdout":
raise RuntimeError(
f"the holdout backtest registered under stage={registered_stage!r} rather than "
"'holdout'; the split-based inference in registry.store._infer_stage did not take "
"precedence over this configuration's risk block"
)
```
## 4. What it came out at
The two numbers below are one strategy measured on two disjoint periods, and the gap between
them is not an estimate of decay. The validation figure is the maximum of a ranking over more
than a thousand backtests, so it carries the selection; the holdout figure is one measurement
over a much shorter window, so it carries that window's sampling error. Both push the pair
apart on their own, before any real change in the strategy's edge.
[`20_strategy_analysis`](20_strategy_analysis.ipynb) is where they are given intervals and a
paired comparison.
```python
metrics = result.metrics
# The selected configuration's own registered Sharpe, not the resolver's. `resolve_solvent_carrier`
# reports the common-support figure, which re-ranks candidates on the timestamps every one of them
# covers; that is the right number for choosing between candidates and the wrong one to set beside a
# holdout measured over its own full window. Both are printed, so neither has to be inferred from
# the other.
with sqlite3.connect(str(CASE_DIR / "run_log" / "registry.db")) as conn:
carrier_sharpe, carrier_periods, carrier_trades = conn.execute(
"SELECT sharpe, n_periods, num_trades FROM backtest_metrics WHERE backtest_hash = ?",
(carrier["val_backtest_hash"],),
).fetchone()
# `n_periods` is on the DAILY grid, not the decision grid. The backtester aggregates to daily
# returns before it computes anything, and `evaluation.periods_per_year` (252) annualizes that
# grid - which is why a six-month intraday holdout reports a few hundred periods rather than
# tens of thousands. Reading it as decision slots is the error that made the committed
# NASDAQ-100 benchmark understate itself roughly fivefold before #362 regenerated it.
print(f"Validation Sharpe over {int(carrier_periods):,} sessions: {carrier_sharpe:.3f}")
print(f" the same run re-ranked on common support: {carrier['val_sharpe']:.3f}")
print(
f"Holdout Sharpe over {int(metrics['n_periods']):,} sessions: "
f"{metrics.get('sharpe', float('nan')):.3f}"
)
print(
f"Holdout: CAGR {metrics.get('cagr', float('nan')):.1%}, "
f"max drawdown {metrics.get('max_drawdown', float('nan')):.2%}, "
f"win rate {metrics.get('win_rate', float('nan')):.0%}"
)
# This case study runs the bar-by-bar engine - the configuration it overlays declares a trailing
# stop, which a vectorized weight-times-return path cannot express - so trade counts are recorded
# and can be compared. A holdout that rebalanced far less than the validation run at the same
# cadence would say the basket stopped changing, which is a different thing from a lower Sharpe.
print(
f"Trades: {int(metrics.get('num_trades', 0)):,} on the holdout, "
f"{int(carrier_trades):,} on validation"
)
```
## What this notebook establishes, and what it does not
It establishes a return series for the selected configuration over a period no choice in this
case study was made on. That is the only thing a holdout can give, and it is worth less than
it looks: the window is short beside the validation span it is being compared with, which is
too few observations to separate a strategy that decayed from one that had an ordinary year.
It does not establish that this configuration was the right one to carry here. The selection
that brought it was made on validation, over a pool large enough that its maximum is
optimistic by construction, and this notebook inherits that pool without correcting for it.
The deflation is [`20_strategy_analysis`](20_strategy_analysis.ipynb)'s.
Re-running this notebook against the same configuration is free and idempotent - the backtest hash
is unchanged and the registered run is served back. A different configuration is refused, for the
reason 18 gives.
**Next:** [`20_strategy_analysis`](20_strategy_analysis.ipynb).Vollständig mit Quellenangabe unter der Lizenz der Quelle angezeigt. Lizenz: MIT
Diese Zusammenfassung wurde vom Research-Agenten von Stratmill anhand des Originals verfasst; sie ist keine Kopie der Quelle.