홀드아웃 누수 없이 CME 선물 전략 선택 및 검증
노트북 Machine Learning for Trading
요약
이 노트북은 신호 규칙, 자산 배분, 리스크 오버레이를 다루는 등록된 검증 백테스트에서 CME 선물 사례 연구를 선택합니다. 두 수익률 구간에 걸쳐 후보를 비교하고 검증 샤프가 가장 높은 설정을 선택합니다. 비용 민감도는 선택 후보군에 후보를 추가하지 않고 선택된 설정만 설명합니다. 재현 가능한 비교를 위해 기준 실행에서 후보 구성을 고정합니다.
그런 다음 선택된 전략의 정확한 전략 사양과 모델 체크포인트를 맞춰 이후 홀드아웃 기간에서 평가합니다. 검증과 결과가 다르더라도 홀드아웃 결과를 보고하며, 이를 근거로 새로 선택하지 않습니다. 코호트의 디플레이티드 샤프 지표와 짝지은 수익률 비교도 갱신하며, 검증 기간과 홀드아웃 기간의 비교도 포함합니다. 각 단계가 이전 단계의 승자를 유지하므로 단계별 중앙값을 같은 모집단에서 나온 것처럼 비교할 수 없다고 주의를 줍니다. 제공된 발췌에는 전체 수치 결과가 없으며, 관련 평가가 실행되지 않았다면 홀드아웃 행이 없을 수 있습니다.
핵심 아이디어
- 신호, 자산 배분, 리스크 오버레이 후보 전반의 검증 샤프를 기준으로 사례 연구를 선택하세요.
- 수익률 구간을 미리 고정하지 말고 나머지 설정과 함께 선택하세요.
- 비용 민감도 실행은 설명용으로 유지해 선택 후보군을 늘리지 마세요.
- 이미 선택된 전략 사양에 맞춰 홀드아웃 재생 결과를 확인하고, 전략을 다시 선택하지 않은 채 보고하세요.
- 후속 단계에는 검증을 거쳐 선택된 생존 후보가 포함되므로 단계별 중앙값은 신중히 해석하세요.
태그
전문
# CME Futures: Strategy Analysis
# CME Futures: Strategy Analysis
The four preceding notebooks each produced a registered, complete population of validation
backtests: the equal-weight signal baseline over every model configuration and checkpoint, the
alternative allocators over the signal shortlist, the transaction-cost grid, and the position-risk
overlays. This notebook reads those registered results and names one case-study configuration
from them.
A case-study configuration is a model configuration - family, settings, and the training
checkpoint the predictions came from - together with the signal, sizing, and risk rules applied to
it. The configuration with the highest validation Sharpe across the signal, allocation, and
risk-overlay stages is the one selected. Cost-sensitivity backtests vary the friction assumption
on a configuration already chosen, so they describe it rather than compete with it and are not in
the pool.
Both return horizons are selected from together. The horizon is part of the configuration, so it
is read off the row that wins rather than fixed before looking.
The holdout period is a later date range that no notebook up to this point has touched. It is
evaluated on the selected configuration alone, one time, and it may disagree with the validation
result. That disagreement is an outcome to report, not a reason to select again.
Prerequisites: `13_backtest`, `14_portfolio_management`, `15_risk_management`, and `16_costs`.
```python
"""Select and describe one CME futures case-study configuration."""
import plotly.express as px
import polars as pl
from case_studies.cme_futures.research_workflow import (
ALL_LABELS,
final_selection_candidate_set,
final_validation_candidate_set,
final_validation_results,
open_study,
product_universe_table,
selection_catalog,
)
from case_studies.research import OfficialPopulation, Result
from case_studies.utils.cohort_metrics import compute_and_register
from case_studies.utils.paired_metrics import populate_paired_metrics
from case_studies.utils.strategy_analysis import (
resolve_solvent_carrier,
select_holdout_self_backtest,
)
from case_studies.utils.uncertainty import ENTIRE_REGISTRY
from utils.style import COLORS
```
```python
EXECUTION_TIER = "canonical"
WORKSPACE: str | None = None
PREVIEW_LABELS: list[str] = []
# The per-label candidate sets this notebook freezes are immutable under their names too, and
# for the same reason as the population above: `CandidateSet.create` refuses a changed member
# list under a name that already exists. Nothing reached that argument before, so any run whose
# membership moved - which a wider sweep does by construction - stopped at the freeze after the
# fit, with no parameter able to answer it.
#
# Each name maps to the generation this run retires. `"live"` names the lineage and looks the
# generation up, which is the form that does not decay: naming the head instead is correct only
# until the next publish, because `create` accepts the head and nothing else. The declaration is
# resolved through `candidate_set_supersedes` rather than offered straight, so a reader's clean
# clone - which has no generation to replace, and often no `candidate_sets` table at all -
# publishes generation one instead of being refused. An unchanged re-run never reads it: a set's
# hash is computed from its members and its contract, so the existing name binding answers.
SUPERSEDES_CANDIDATE_SETS: dict[str, str] = {
"cme_futures-pre-overlay-fwd_ret_5d-v1": "live",
"cme_futures-pre-overlay-fwd_ret_21d-v1": "live",
"cme_futures-final-validation-fwd_ret_5d-v1": "live",
"cme_futures-final-validation-fwd_ret_21d-v1": "live",
"cme_futures-final-selection-v1": "live",
}
```
## The pool the configuration is selected from
Each per-label pool opens the three stage populations and fails if any member is missing,
incomplete, or produced under a preview identity. Combining the two horizons into one immutable
set records exactly which results were compared, so the selection can be repeated later against
the same members rather than against whatever the registry holds at the time.
```python
study = open_study(execution_tier=EXECUTION_TIER, workspace=WORKSPACE)
if EXECUTION_TIER == "canonical":
if PREVIEW_LABELS:
raise ValueError("canonical execution cannot declare preview reductions")
labels = ALL_LABELS
elif EXECUTION_TIER == "preview":
if WORKSPACE is None or not PREVIEW_LABELS:
raise ValueError("preview execution requires WORKSPACE and PREVIEW_LABELS")
unknown = sorted(set(PREVIEW_LABELS) - set(ALL_LABELS))
if unknown:
raise ValueError(f"preview labels this case study does not declare: {unknown}")
labels = tuple(PREVIEW_LABELS)
else:
raise ValueError(f"unsupported execution tier: {EXECUTION_TIER!r}")
universe = product_universe_table()
universe
```
Only a canonical pool is an immutable set. A preview run publishes no candidate set - one cannot
hold a preview member - so its pool is the rows its own reduced execution produced and the
`candidate_set_hash` column below is null. Everything downstream, the ranking rule included, is
the same either way; what differs is whether the pool can be reopened later by name.
```python
if EXECUTION_TIER == "canonical":
per_label = {
label: final_validation_candidate_set(
study, label=label, supersedes_by_set=SUPERSEDES_CANDIDATE_SETS
)
for label in labels
}
per_label_results = {
label: tuple(Result.open(study, value) for value in pool_set.members)
for label, pool_set in per_label.items()
}
candidates = final_selection_candidate_set(study, supersedes_by_set=SUPERSEDES_CANDIDATE_SETS)
pool_results = tuple(Result.open(study, value) for value in candidates.members)
pool_identity = candidates.hash
per_label_identity = {label: pool_set.hash for label, pool_set in per_label.items()}
else:
per_label_results = {
label: final_validation_results(study, label=label, execution_tier=EXECUTION_TIER)
for label in labels
}
pool_results = tuple(result for results in per_label_results.values() for result in results)
pool_identity = None
per_label_identity = dict.fromkeys(labels)
pool = selection_catalog(study, (result.hash for result in pool_results))
pool_size = pl.DataFrame(
[
{
"label": label,
"candidates": len(results),
"candidate_set_hash": per_label_identity[label],
}
for label, results in per_label_results.items()
],
schema={"label": pl.String, "candidates": pl.Int64, "candidate_set_hash": pl.String},
).sort("label")
```
## The selection correction, and the paired comparisons
Two registry tables carry the statistics this notebook reports rather than recomputes:
`cohort_metrics` holds the deflated Sharpe for each cohort's leader, and
`backtest_paired_metrics` holds bootstrap comparisons between registered return series.
Both were empty here until 2026-08-31, so an earlier edition of the README quoted deflation
numbers with nothing behind them. The cause was not a bad computation - it was that this
notebook never called for one, while `etfs`, `fx_pairs` and `us_firm_characteristics` all do.
`compute_and_register` refreshes the whole table rather than one row, so it can never report a
stale leader. `populate_paired_metrics` writes one row per comparison kind, including
`val_rank1_self` - the selected configuration's validation series against its own holdout replay,
which is the paired form of the val-to-holdout question and the only honest way to ask it.
Comparing two point estimates is not that question: the holdout is a shorter window, so the
difference carries sampling error the point estimates do not show.
The selected configuration is resolved here rather than further down because
`populate_paired_metrics` needs it. Omitting it does not fail - it falls back to ranking the
registry on raw Sharpe, which on this registry names `latent_factors`/`sdf` on `fwd_ret_21d`,
while the canonical resolver names `gbm`/`leaves_31_mse` on `fwd_ret_5d`. The paired rows would
then compare a strategy the chapter does not report, under headings that say they describe the
one it does. That is the same disagreement documented below for the holdout lookup, reaching a
different table.
`replace_all=True` makes the call a snapshot rather than an insert. Registration is an upsert
keyed on the pair, so it cannot remove rows a previous selection wrote; without the prune, the
raw-Sharpe pairs would survive alongside the selected configuration's.
`prediction_hashes` scopes the cohorts to this notebook's own pool. On this registry it changes
nothing - the cohorts are already a strict subset of the pool, because it was rebuilt from empty
and holds no retired generation. That is a property of the registry, not of the call: without the
argument, a superseded generation left in the registry would inflate K and could lead a cohort
outright, and the deflation this notebook publishes would be computed over a variant the pool
excludes. Being right by accident is not the same as being right.
```python
carrier = resolve_solvent_carrier("cme_futures")
cohort_counts = compute_and_register(
"cme_futures",
prediction_hashes=pool.get_column("prediction_hash").unique().to_list(),
verbose=False,
)
# The cohort call above is scoped to the reported pool and this one is not: the pairs are
# selected from every registered prediction set. Stated rather than defaulted; narrowing
# it would change published numbers, so it is a separate decision from this line.
paired_rows = populate_paired_metrics(
"cme_futures",
carrier=carrier,
replace_all=True,
prediction_hashes=ENTIRE_REGISTRY,
verbose=False,
)
print(f"cohort_metrics: {sum(cohort_counts[k] for k in ('family', 'stagelabel', 'label'))} rows")
print(f"backtest_paired_metrics: {sum(1 for r in paired_rows if 'skip' not in r)} pairs")
```
```python
pool_size
```
## What each selection stage contributed
The three stages run in sequence, each on the survivors of the one before. The baseline stage
carries every configuration and checkpoint at equal weight. Allocation runs on the strongest
distinct configurations from that stage, and the risk overlay on the strongest result so far for
each horizon. Later stages therefore hold far fewer candidates than the first, and the spread
within a stage shows how much of the outcome the sizing and risk rules decide once the model is
fixed.
**The medians cannot be read across stages.** Each stage runs on the survivors of the one before,
chosen on the same validation Sharpe the table reports, so the pool shrinks from 496 to 60 to 14
by selecting on the quantity being summarized. On `fwd_ret_21d` the median rises from -0.392 to
0.192 to 1.010 along that shrinking pool, and almost all of that movement is the selection, not
the position sizing methods or the risk rules. What the stages do support is the comparison
within a row: the fourteen risk overlays share one model, one signal and one sizing rule,
and they still span 0.322 to 1.274, which is the range the position rule alone is
responsible for.
```python
stage_summary = (
pool.group_by("label", "stage")
.agg(
pl.len().alias("candidates"),
pl.col("sharpe").min().alias("min_sharpe"),
pl.col("sharpe").median().alias("median_sharpe"),
pl.col("sharpe").max().alias("max_sharpe"),
)
.sort("label", "stage")
)
```
```python
stage_summary
```
```python
fig = px.strip(
pool.to_pandas(),
x="stage",
y="sharpe",
color="label",
stripmode="overlay",
category_orders={"stage": ["signal", "allocation", "risk_overlay"]},
color_discrete_sequence=[COLORS["blue"], COLORS["amber"]],
labels={"stage": "Selection stage", "sharpe": "Validation Sharpe", "label": "Return horizon"},
)
fig.update_layout(
title="Validation Sharpe by selection stage and return horizon",
height=420,
)
fig.add_hline(y=0, line_dash="dash", line_color=COLORS["neutral"])
fig.show()
```
## The selected configuration
The selection is made by `resolve_solvent_carrier`, the shared resolver, and not by ranking this
pool's Sharpe column directly. The two do not agree here. Ranking the column names the
`latent_factors` / `sdf` row on `fwd_ret_21d` at 1.274; the resolver names the `gbm` /
`leaves_31_mse` row on `fwd_ret_5d`, whose raw 1.236 becomes 1.294 once the candidates are
compared over the 1,270 sessions they all price. Different family, different horizon, from the
same registry.
The re-ranking is the reason to prefer the resolver. A Sharpe computed over a configuration's own
available history is not comparable across configurations that priced different spans, and
ranking the raw column silently rewards whichever candidate had the most forgiving window. The
resolver also refuses a selected configuration that is insolvent rather than reporting it.
It matters here beyond correctness of the ranking. `17_holdout_predictions` and
`18_holdout_backtest` resolve it the same way, so a second selection rule
in this notebook would ask `select_holdout_self_backtest` for the holdout replay of a
configuration those notebooks never ran. The answer would be `None`, and this notebook would
report the holdout as not produced while it sat in the registry.
The prediction checkpoint is part of the identity either way: two rows from the same trained
model at different checkpoints are different configurations, and a holdout matched on the
trained model alone can land on a different checkpoint from the one selected.
```python
selected = next(
(result for result in pool_results if result.hash == carrier["val_backtest_hash"]), None
)
if selected is None:
raise RuntimeError(
f"the resolved configuration {carrier['val_backtest_hash']} ({carrier['family']}/"
f"{carrier['config_name']}, {carrier['label']}, stage {carrier['val_stage']}) is not in "
"this notebook's pool. The pool and the shared resolver are reading the same registry, so "
"they disagree about which stages are selected from, and the holdout notebooks followed "
"the resolver."
)
selected_row = pool.filter(pl.col("backtest_hash") == selected.hash)
selected_label = selected_row.item(0, "label")
selected_strategy = selected.spec()["strategy"]
```
```python
selected_row
```
```python
pl.DataFrame(
[
{
"candidate_set_hash": pool_identity,
"candidates_compared": len(pool_results),
"label": selected_label,
"signal": str(selected_strategy["signal"]),
"allocation": str(selected_strategy.get("allocation")),
"risk": str(selected_strategy.get("risk")),
}
]
)
```
## What friction costs this configuration
The cost grid was run on the single configuration this case study ships - the same one selected
above, resolved across labels and priced with its risk overlay in place - holding the model,
sizing, risk rules and contract specification fixed and varying only the all-in cost assumption.
Commission and slippage each take half of the quoted figure. One curve, not one per horizon:
there is one strategy, so the label the selected configuration does not sit on has no cost rows
at all.
```python
if EXECUTION_TIER == "canonical":
cost_population = OfficialPopulation.one(study, name="cme_futures-cost-validation-v1")
cost_members = list(cost_population.require_complete())
else:
cost_members = (
study.backtests.table(include_preview=True)
.filter(
(pl.col("execution_tier") == "preview")
& (pl.col("stage") == "cost_sensitivity")
& pl.col("complete")
)
.get_column("backtest_hash")
.to_list()
)
cost_curve = (
study.backtests.table(include_preview=True)
.filter(pl.col("backtest_hash").is_in(cost_members) & (pl.col("label") == selected_label))
.with_columns(
(
pl.col("spec_json")
.str.json_path_match("$.decision_artifact.parameters.costs.commission_bps")
.cast(pl.Float64)
+ pl.col("spec_json")
.str.json_path_match("$.decision_artifact.parameters.costs.slippage_bps")
.cast(pl.Float64)
).alias("total_cost_bps")
)
.select("total_cost_bps", "sharpe", "total_return", "num_trades", "backtest_hash")
.sort("total_cost_bps")
)
if cost_curve.is_empty():
raise RuntimeError(f"the cost population contains no member for {selected_label!r}")
# `json_path_match` returns null for a path that is not in the document rather than raising, and
# null + null is null, so reading the grid value from the wrong place yields a full-height frame
# whose cost axis is entirely missing. The emptiness check above passes on such a frame and the
# curve below plots against nothing. Refuse it here instead.
if cost_curve.get_column("total_cost_bps").null_count():
raise RuntimeError(
"cost members record no all-in cost at "
"$.decision_artifact.parameters.costs; the cost curve has no axis"
)
```
```python
cost_curve
```
```python
fig = px.line(
cost_curve.to_pandas(),
x="total_cost_bps",
y="sharpe",
markers=True,
labels={"total_cost_bps": "All-in cost (bps per trade)", "sharpe": "Validation Sharpe"},
color_discrete_sequence=[COLORS["copper"]],
)
fig.update_layout(
title="Validation Sharpe across the all-in transaction-cost grid",
height=380,
)
fig.add_hline(y=0, line_dash="dash", line_color=COLORS["neutral"])
fig.show()
```
## The holdout
The holdout evaluates one configuration: the one the validation backtests selected. That is the
highest validation backtest Sharpe across the baseline, position-sizing, allocation and
risk-management stages, and it is fixed before any holdout artifact exists.
What keeps the holdout from becoming an axis to search over is the direction of that rule, not a
gate. The ranking reads validation rows only, and the holdout row below is found by matching the
selected strategy specification - never by taking whichever holdout backtest scored best. A
holdout number therefore cannot change which configuration is reported here.
Nothing about it is one-shot. A holdout result that turns out to be wrong is deleted and produced
again; what would make the number uninterpretable is evaluating many configurations on the window
and reporting the best, which is the thing the selection rule rules out. The holdout notebooks
produce the row; this one reads it. Where they have not run, the table is empty and the validation
result above stands on its own.
```python
# `select_holdout_self_backtest` is the shared resolver every strategy-analysis notebook uses.
# It takes the selection this notebook already made and finds the holdout backtest replaying that
# same strategy specification, at the same configuration and checkpoint, over a training run whose
# own CV declares the holdout fold. It returns None where no such run exists, and raises rather
# than choosing where two of them do.
#
# Calling it rather than re-deriving the lineage here is deliberate. A second implementation
# living beside the first agrees with it on the registry it was written against and diverges on
# the next one, and a divergence in this particular lookup is a holdout number attributed to the
# wrong configuration.
holdout_backtest_hash = select_holdout_self_backtest("cme_futures", selected.hash)
print(
f"Selected validation backtest: {selected.hash} ({selected_label})\n"
f"Holdout replay: {holdout_backtest_hash or 'not produced yet'}"
)
```
```python
if holdout_backtest_hash is None:
comparison = pl.DataFrame()
else:
evaluated = study.backtests.table(include_preview=True).filter(
pl.col("backtest_hash") == holdout_backtest_hash
)
comparison = pl.concat(
[
selected_row.select("label", "sharpe", "max_drawdown", "num_trades").with_columns(
pl.lit("validation").alias("split")
),
evaluated.select("label", "sharpe", "max_drawdown", "num_trades").with_columns(
pl.lit("holdout").alias("split")
),
]
).select("split", "label", "sharpe", "max_drawdown", "num_trades")
comparison
```출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: MIT
이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.