CME 선물 모델의 정보계수 해석
노트북 Machine Learning for Trading
요약
이 분석은 CME 선물 모델의 횡단면 정보계수(IC)를 해석하는 방법과 이후 포트폴리오 테스트와의 관계를 설명합니다. 각 날짜의 예측 수익률과 실제 수익률 간 순위 상관관계를 계산한 다음, 날짜별 상관계수의 평균을 구합니다. 이는 예측 수익률의 크기는 무시하고 순위에 따라 상품을 선택하는 동일 가중 전략에 해당합니다. 또한 폴드 간 평균과 t 통계량을 일별 시계열 추론과 구분하며, 보고된 t 통계량은 폴드 간 표준오차를 사용하고 예측값이 거의 일정해질 때는 IC가 정의된 날짜 수가 중요하다고 설명합니다.
이 분석은 IC를 선택 규칙이 아니라 진단 지표로 다룹니다. 회전율 비용이나 제한된 유동성이 신호를 압도하면 순위 예측력이 높아도 수익성 있는 포트폴리오를 만들지 못할 수 있습니다. 완성된 모든 모델 설정은 검증 백테스트로 넘어가며, 샤프 비율을 기준으로 전략과 체크포인트를 선택합니다. 노트북은 인과 분석의 반박 열을 위약 검정의 견고성을 보여주는 신뢰할 만한 증거로 보기 어렵다고 설명하고, 대신 DML 추정치와 HAC 불확실성을 참고하도록 안내합니다. 결론은 예측 수준의 지표로 한정되며, 포트폴리오 구성, 자금 조달 및 체결은 다른 곳에서 평가합니다.
핵심 아이디어
- IC는 예측 크기의 정확도가 아니라 날짜별 예측 수익률과 실제 수익률 간 순위 연관성을 측정합니다.
- 공통 시장 움직임이 횡단면 예측력을 왜곡하지 않도록 각 의사결정 날짜에 IC를 계산하고 날짜별 평균을 구합니다.
- 보고된 IC t 통계량은 폴드 수준 표준오차를 사용하므로 평균과 함께 날짜별 데이터 범위도 살펴봐야 합니다.
- 회전율과 유동성이 순위의 거래 가능성에 영향을 주므로 높은 IC만으로 수익성이 입증되지는 않습니다.
- 설정과 체크포인트는 IC가 아니라 백테스트 포트폴리오 성과로 선택합니다.
태그
전문
# CME Futures: Model Analysis
# CME Futures: Model Analysis
This notebook reads the complete canonical model population produced by `06_linear` through
`10b_stochastic_discount_factor`. Each row retains family, configuration, label, checkpoint, fold
contract, training identity, and prediction identity. No null label is assigned to another
horizon, and checkpoints are not collapsed into a single model label.
IC measures whether a configuration ranks the cross-section correctly on a decision date. It
selects nothing: every row here proceeds to the equal-weight validation backtest in
`13_backtest`, where Sharpe performs selection and the checkpoint is part of what is selected.
What `ic_mean` and `ic_t` are, precisely, because the two readings are easy to confuse. Both are
computed **across folds**: `ic_mean` averages each fold's cross-sectional IC, and `ic_t` divides
that by the *standard error* of the same five numbers, which is their dispersion divided by the
square root of how many are defined - not the dispersion itself. `ic_std` in the table below is
the dispersion, so reproducing `ic_t` from the two columns needs the `sqrt(n_folds_ic)` factor.
The registry also computes the daily-series reading with its HAC standard error, which is the
inferential statistic, but the predictions reader does not surface it, so it is not in the table
below. `ic_n_days` is carried instead: it
counts the validation dates that produced a defined IC, and a configuration whose predictions
collapse to near-constant on some dates has its `ic_mean` measured over fewer of them. Reading
`ic_mean` without `ic_n_days` is how a partial-coverage artifact reads as a leader.
## What an information coefficient measures, and what it cannot
Everything in the table below is built on the IC, so it is worth being exact about what the
number is before reading any of it.
On one decision date there is a set of products, a predicted return for each, and the return
each actually went on to earn. The IC is the rank correlation between those two lists. It asks
a deliberately narrow question: did the model put them **in the right order**? It does not ask
whether the predicted magnitudes were close, and it cannot - a model that predicts every return
at a hundredth of its true size scores exactly as well as one that gets the levels right, as
long as the ordering matches.
That narrowness is the point for a strategy of this shape. The backtest in `13_backtest` holds
the top-ranked products and shorts the bottom-ranked ones in equal weight, so the ordering is
the entire input and the magnitudes are discarded before a position is taken. A diagnostic that
rewarded accurate levels would be measuring something the strategy never uses.
**It is computed per decision date and then averaged, never pooled.** Pooling every
product-date into one correlation would let a period when the whole market moved together
masquerade as skill at telling products apart: on a day when everything rallies, a model that
ranks products at random still shows agreement between its predictions and the outcomes if the
cross-sections are stacked. Correlating within a date and averaging afterwards removes the
common move by construction, because it is the same for every product on that date.
### Why a good IC is a small number
A reader arriving from a forecasting background should expect these to look disappointing. A
monthly-horizon equity or futures IC of 0.03 to 0.05 is a real, usable signal; 0.10 sustained
would be remarkable. This is not a weak result being excused - it is what predicting an
overwhelmingly noise-dominated quantity looks like when it works. The edge comes from applying
a small consistent tilt across many products and many dates, not from being right about any one
of them, and the arithmetic of that is what `13_backtest` measures and this notebook does not.
### Why the ranking diagnostic selects nothing
A high IC does not imply a tradeable strategy, and this is the boundary the notebook's title
refers to. A configuration can rank the cross-section well and still lose money: if its ranking
churns from one decision to the next, the turnover it implies costs more than the tilt earns; if
its skill sits entirely in the products that are least liquid, the positions cannot be taken at
the prices the backtest assumes. Neither of those is visible in an IC, because an IC has no
notion of holding anything.
So nothing here is a decision. Every complete row in this catalog proceeds to the equal-weight
validation backtest, where Sharpe selects and the checkpoint is part of what is selected. This
notebook is where a reader forms an expectation and, more usefully, notices where the backtest
later disagrees with it - a configuration that ranks well and backtests badly is the most
informative row in the whole case study, because the gap between the two is exactly where
turnover and tradeability live.
```python
"""Analyze complete CME futures model and causal result catalogs."""
import json
import numpy as np
import polars as pl
from case_studies.cme_futures.research_workflow import (
ALL_LABELS,
CASE_STUDY,
MODEL_POPULATION_NAMES,
official_prediction_catalog,
open_study,
product_universe_table,
)
from case_studies.research import CausalResult, require_declared_menu_coverage
from utils.paths import get_case_study_dir
```
```python
EXECUTION_TIER = "canonical"
WORKSPACE: str | None = None
```
## Complete prediction catalog
The six official population snapshots were created before their model runs. Opening all six and
calling `require_complete` means a failed configuration or checkpoint cannot disappear from this
analysis because another row happened to finish.
```python
if EXECUTION_TIER == "preview" and WORKSPACE is None:
raise ValueError("preview execution requires WORKSPACE")
study = open_study(execution_tier=EXECUTION_TIER, workspace=WORKSPACE)
universe = product_universe_table()
universe
```
Canonical analysis reads the six published population snapshots, which is the whole point of
freezing them before their runs. A preview has no published population to read: it is a reduced
re-execution whose rows exist only in its own workspace and which is deliberately excluded from
every official population. It reads its own complete validation predictions instead, and the
comparison against the declared menus below is skipped with them, because a preview fits a named
subset by design and would fail that comparison on every configuration it left out.
```python
if EXECUTION_TIER == "canonical":
catalog = official_prediction_catalog(study, MODEL_POPULATION_NAMES)
else:
catalog = (
study.predictions.table(include_preview=True)
.filter(
(pl.col("execution_tier") == "preview")
& (pl.col("split") == "validation")
& pl.col("complete")
)
.sort("label", "family", "config_name", "checkpoint_kind", "checkpoint_value")
)
if catalog.is_empty():
raise RuntimeError("preview execution registered no complete validation predictions")
def _feature_count(spec_json: str) -> int:
spec = json.loads(spec_json)
computation = spec.get("computation", spec)
return len(computation.get("feature_names") or [])
analysis = catalog.with_columns(
pl.col("spec_json").map_elements(_feature_count, return_dtype=pl.Int64).alias("feature_count")
).select(
"family",
"config_name",
"label",
"checkpoint_kind",
"checkpoint_value",
"feature_count",
"ic_mean",
"ic_t",
"ic_n_days",
"n_folds",
"training_hash",
"prediction_hash",
)
```
### Every declared model is here
Each execution notebook checks that it produced everything **it** requested, so none of them can
see a configuration that no notebook requests at all - a menu entry nobody claimed publishes
nothing and every completeness check still passes. This is the one place the families reassemble,
so it is the only place that check can be made. `require_declared_menu_coverage` compares
`(family, label, config_name)` against the training menus and raises on either direction: a
declared model the population omits, or a model in the population that no menu declares.
It returns the rows knowingly excluded, so what this notebook is missing is displayed rather than
taken on trust. `causal_dml` is not in the comparison - it is not a predictive family and the
adapter registry, not a list here, is what decides that.
```python
if EXECUTION_TIER == "canonical":
excluded = require_declared_menu_coverage(analysis, case_study=CASE_STUDY)
else:
excluded = analysis.clear()
excluded
```
```python
analysis.sort("label", "family", "config_name", "checkpoint_value")
```
## Interpretation boundaries
The table compares ranking diagnostics under the declared walk-forward protocol. The backtest
engine supplies portfolio returns, transaction costs, contract sizing, and roll execution before
selection.
The distinction is worth stating as a rule rather than as a caveat: **every quantity in this
notebook is a property of the predictions, and every quantity that decides anything is a
property of a portfolio.** A prediction has no size, no holding period, no financing and no
execution price. Those enter in `13_backtest`, and they are capable of reordering the table
below completely - which is why a leaderboard here is a hypothesis about the backtest rather
than a preview of it.
Conformal weighting, when used later as an allocator, calibrates chronologically from prior
validation observations. It uses the calibration-window scale and the finite-sample higher order
statistic. It does not fit a scale on the evaluation fold or pool all folds before calibration.
## Causal diagnostics
Double machine learning answers a different question from prediction. The treatment effect is
conditioned on the configured confounders, and HAC uncertainty follows the decision-time order.
A covariance-estimator failure is not relabeled as HAC. The shared runner must return a finite HAC
standard error for a result to be complete.
**The `refutation_p` column below is not evidence that these effects survived a placebo test.**
The refutation permutes contiguous blocks within each product, and the shared runner sizes those
blocks as `max(label_buffer, treatment_window)`. `causal.treatment_window` is 1 here, so the label
buffer binds and the registered rows carry a 21-period block for `fwd_ret_21d` and a 5-period
block for `fwd_ret_5d`. Neither length is a property of `carry_pct`, whose own persistence the
cell below measures on this case study's feature panel: the autocorrelation is pooled within
product, each product demeaned before pooling so a level difference between products cannot
stand in for persistence within one, and on one row per product-session, because the block
counts sessions.
**The two blocks sit at very different points on that profile, so the concern bears much more
on one label than the other.** Read the block lengths against the autocorrelation at those
lags rather than against the half-life: the decay is slower than the AR(1) half-life implies -
an AR(1) with this lag-1 value would sit at 0.38 by lag 5 and 0.02 by lag 21, where the panel
is at 0.52 and 0.14 - so the half-life is a lower bound on persistence, not the yardstick for
the block. At the 5-session block used for `fwd_ret_5d` the autocorrelation is still 0.52, so that
block leaves real dependence unpreserved; the placebo is a weaker opponent than the truth and
`fwd_ret_5d`'s empirical p-value is biased toward zero by some amount this notebook does not
quantify. At the 21-session block used for `fwd_ret_21d` it is 0.14, and indistinguishable
from zero by lag 63, so that block spans most of the dependence and the concern is
correspondingly weaker there.
**That is no longer what the column reports, and the reason is worth following.** `fwd_ret_5d`
used to sit at 0.0396 and `fwd_ret_21d` at 0.0099, which is 1/101 and the floor 100 draws can
report. The refutation now compares HAC t-statistics rather than raw effects, because a permuted
treatment is not predictable from the controls, its residual keeps nearly all its variance, and
that variance is the denominator of the second-stage effect - so every placebo effect was divided
by a larger number than the observed one. Correcting that moved `fwd_ret_5d` to 0.5545 and
`fwd_ret_21d` to 0.2673, both `Fails`, on an identical fit. The block-length argument above is a
separate, uncorrected narrowing and it bears mainly on `fwd_ret_5d`; either way it is no longer
visible in these two numbers. Read the DML point estimate and its HAC standard error. The
refutation column is recorded for completeness and carries no evidence here.
```python
# The panel carries one row per contract position, so a product-session appears up to three
# times with the same carry_pct. Lagging without de-duplicating steps ~2.78 rows per session
# and reports a persistence profile stretched by that factor. The block the refutation permutes
# counts sessions - run_dml_analysis requires strictly increasing timestamps within a product -
# so sessions are the scale the two have to be compared on.
carry = (
pl.read_parquet(get_case_study_dir(CASE_STUDY) / "features" / "financial.parquet")
.select(["product", "timestamp", "carry_pct"])
.drop_nulls()
.unique(subset=["product", "timestamp"])
.sort(["product", "timestamp"])
)
autocorr = []
for lag in (1, 5, 21, 63):
paired = (
carry.with_columns(pl.col("carry_pct").shift(lag).over("product").alias("lagged"))
.drop_nulls()
.with_columns(
(pl.col("carry_pct") - pl.col("carry_pct").mean().over("product")).alias("x"),
(pl.col("lagged") - pl.col("lagged").mean().over("product")).alias("y"),
)
)
rho = (paired["x"] * paired["y"]).sum() / (
((paired["x"] ** 2).sum() * (paired["y"] ** 2).sum()) ** 0.5
)
autocorr.append({"lag": lag, "autocorrelation": rho, "n_pairs": paired.height})
carry_persistence = pl.DataFrame(autocorr)
half_life = np.log(0.5) / np.log(carry_persistence["autocorrelation"][0])
print(
f"carry_pct within-product pooled autocorrelation, "
f"AR(1) half-life {half_life:.1f} sessions from lag 1"
)
carry_persistence
```
```python
causal_rows = []
for label in ALL_LABELS:
result = CausalResult.one(study, label=label, execution_tier=EXECUTION_TIER)
if not result.complete:
raise RuntimeError(f"causal result for {label} is incomplete")
causal_rows.append({"label": label, "causal_hash": result.hash, **result.metrics})
causal = pl.DataFrame(causal_rows).sort("label")
```
```python
causal
```
## What proceeds to backtesting
All complete prediction rows proceed. The next notebook passes the selected Polars rows directly
to the shared backtest call, publishes product-keyed decisions, and records contract, roll, price,
and prediction lineage. Validation backtest Sharpe, with the prediction checkpoint included in the
configuration identity, is the selection statistic.출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: MIT
이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.