Interpréter les coefficients d’information des modèles de contrats à terme CME
Résumé
Cette analyse explique comment lire les coefficients d’information transversaux (IC) des modèles de contrats à terme CME et les relier aux tests de portefeuille ultérieurs. Elle calcule la corrélation de rang entre rendements prévus et réalisés pour chaque date, puis en fait la moyenne sur l’ensemble des dates. Cela correspond à une stratégie à pondération égale qui sélectionne les produits par rang et ignore l’amplitude des rendements prévus. Le document distingue également la moyenne et la statistique t entre plis de l’inférence sur séries quotidiennes : la statistique t présentée utilise l’erreur standard entre plis, et le nombre de dates avec des IC définis compte lorsque les prévisions deviennent presque constantes.
L’analyse traite l’IC comme un diagnostic, et non comme une règle de sélection. Une forte capacité de classement peut ne pas produire de portefeuille rentable si les coûts de rotation ou la faible liquidité l’emportent sur le signal. Toutes les configurations de modèle complètes passent aux backtests de validation, où le ratio de Sharpe sert à sélectionner les stratégies et les points de contrôle. Le carnet indique que la colonne de réfutation de l’analyse causale n’est pas un indicateur fiable de robustesse aux placebos et renvoie plutôt à l’estimation DML et à l’incertitude HAC. Ses conclusions se limitent aux mesures au niveau des prévisions ; la construction du portefeuille, le financement et l’exécution sont évalués ailleurs.
Idées clés
- IC mesure l’association de rang entre rendements prévus et réalisés à une date donnée, pas la précision de l’amplitude des prévisions.
- Calculez l’IC à chaque date de décision et faites la moyenne sur les dates afin que les mouvements communs du marché ne faussent pas la compétence transversale.
- La statistique t de l’IC présentée utilise l’erreur standard au niveau des plis ; il faut aussi examiner la couverture des dates.
- Un IC élevé n’établit pas la rentabilité, car la rotation et la liquidité influent sur la possibilité de négocier les classements.
- Ce sont les performances des portefeuilles en backtest, et non l’IC, qui déterminent les configurations et les points de contrôle retenus.
Étiquettes
Texte intégral
# CME Futures: Model Analysis
# CME Futures: Model Analysis
This notebook reads the complete canonical model population produced by `06_linear` through
`10b_stochastic_discount_factor`. Each row retains family, configuration, label, checkpoint, fold
contract, training identity, and prediction identity. No null label is assigned to another
horizon, and checkpoints are not collapsed into a single model label.
IC measures whether a configuration ranks the cross-section correctly on a decision date. It
selects nothing: every row here proceeds to the equal-weight validation backtest in
`13_backtest`, where Sharpe performs selection and the checkpoint is part of what is selected.
What `ic_mean` and `ic_t` are, precisely, because the two readings are easy to confuse. Both are
computed **across folds**: `ic_mean` averages each fold's cross-sectional IC, and `ic_t` divides
that by the *standard error* of the same five numbers, which is their dispersion divided by the
square root of how many are defined - not the dispersion itself. `ic_std` in the table below is
the dispersion, so reproducing `ic_t` from the two columns needs the `sqrt(n_folds_ic)` factor.
The registry also computes the daily-series reading with its HAC standard error, which is the
inferential statistic, but the predictions reader does not surface it, so it is not in the table
below. `ic_n_days` is carried instead: it
counts the validation dates that produced a defined IC, and a configuration whose predictions
collapse to near-constant on some dates has its `ic_mean` measured over fewer of them. Reading
`ic_mean` without `ic_n_days` is how a partial-coverage artifact reads as a leader.
## What an information coefficient measures, and what it cannot
Everything in the table below is built on the IC, so it is worth being exact about what the
number is before reading any of it.
On one decision date there is a set of products, a predicted return for each, and the return
each actually went on to earn. The IC is the rank correlation between those two lists. It asks
a deliberately narrow question: did the model put them **in the right order**? It does not ask
whether the predicted magnitudes were close, and it cannot - a model that predicts every return
at a hundredth of its true size scores exactly as well as one that gets the levels right, as
long as the ordering matches.
That narrowness is the point for a strategy of this shape. The backtest in `13_backtest` holds
the top-ranked products and shorts the bottom-ranked ones in equal weight, so the ordering is
the entire input and the magnitudes are discarded before a position is taken. A diagnostic that
rewarded accurate levels would be measuring something the strategy never uses.
**It is computed per decision date and then averaged, never pooled.** Pooling every
product-date into one correlation would let a period when the whole market moved together
masquerade as skill at telling products apart: on a day when everything rallies, a model that
ranks products at random still shows agreement between its predictions and the outcomes if the
cross-sections are stacked. Correlating within a date and averaging afterwards removes the
common move by construction, because it is the same for every product on that date.
### Why a good IC is a small number
A reader arriving from a forecasting background should expect these to look disappointing. A
monthly-horizon equity or futures IC of 0.03 to 0.05 is a real, usable signal; 0.10 sustained
would be remarkable. This is not a weak result being excused - it is what predicting an
overwhelmingly noise-dominated quantity looks like when it works. The edge comes from applying
a small consistent tilt across many products and many dates, not from being right about any one
of them, and the arithmetic of that is what `13_backtest` measures and this notebook does not.
### Why the ranking diagnostic selects nothing
A high IC does not imply a tradeable strategy, and this is the boundary the notebook's title
refers to. A configuration can rank the cross-section well and still lose money: if its ranking
churns from one decision to the next, the turnover it implies costs more than the tilt earns; if
its skill sits entirely in the products that are least liquid, the positions cannot be taken at
the prices the backtest assumes. Neither of those is visible in an IC, because an IC has no
notion of holding anything.
So nothing here is a decision. Every complete row in this catalog proceeds to the equal-weight
validation backtest, where Sharpe selects and the checkpoint is part of what is selected. This
notebook is where a reader forms an expectation and, more usefully, notices where the backtest
later disagrees with it - a configuration that ranks well and backtests badly is the most
informative row in the whole case study, because the gap between the two is exactly where
turnover and tradeability live.
```python
"""Analyze complete CME futures model and causal result catalogs."""
import json
import numpy as np
import polars as pl
from case_studies.cme_futures.research_workflow import (
ALL_LABELS,
CASE_STUDY,
MODEL_POPULATION_NAMES,
official_prediction_catalog,
open_study,
product_universe_table,
)
from case_studies.research import CausalResult, require_declared_menu_coverage
from utils.paths import get_case_study_dir
```
```python
EXECUTION_TIER = "canonical"
WORKSPACE: str | None = None
```
## Complete prediction catalog
The six official population snapshots were created before their model runs. Opening all six and
calling `require_complete` means a failed configuration or checkpoint cannot disappear from this
analysis because another row happened to finish.
```python
if EXECUTION_TIER == "preview" and WORKSPACE is None:
raise ValueError("preview execution requires WORKSPACE")
study = open_study(execution_tier=EXECUTION_TIER, workspace=WORKSPACE)
universe = product_universe_table()
universe
```
Canonical analysis reads the six published population snapshots, which is the whole point of
freezing them before their runs. A preview has no published population to read: it is a reduced
re-execution whose rows exist only in its own workspace and which is deliberately excluded from
every official population. It reads its own complete validation predictions instead, and the
comparison against the declared menus below is skipped with them, because a preview fits a named
subset by design and would fail that comparison on every configuration it left out.
```python
if EXECUTION_TIER == "canonical":
catalog = official_prediction_catalog(study, MODEL_POPULATION_NAMES)
else:
catalog = (
study.predictions.table(include_preview=True)
.filter(
(pl.col("execution_tier") == "preview")
& (pl.col("split") == "validation")
& pl.col("complete")
)
.sort("label", "family", "config_name", "checkpoint_kind", "checkpoint_value")
)
if catalog.is_empty():
raise RuntimeError("preview execution registered no complete validation predictions")
def _feature_count(spec_json: str) -> int:
spec = json.loads(spec_json)
computation = spec.get("computation", spec)
return len(computation.get("feature_names") or [])
analysis = catalog.with_columns(
pl.col("spec_json").map_elements(_feature_count, return_dtype=pl.Int64).alias("feature_count")
).select(
"family",
"config_name",
"label",
"checkpoint_kind",
"checkpoint_value",
"feature_count",
"ic_mean",
"ic_t",
"ic_n_days",
"n_folds",
"training_hash",
"prediction_hash",
)
```
### Every declared model is here
Each execution notebook checks that it produced everything **it** requested, so none of them can
see a configuration that no notebook requests at all - a menu entry nobody claimed publishes
nothing and every completeness check still passes. This is the one place the families reassemble,
so it is the only place that check can be made. `require_declared_menu_coverage` compares
`(family, label, config_name)` against the training menus and raises on either direction: a
declared model the population omits, or a model in the population that no menu declares.
It returns the rows knowingly excluded, so what this notebook is missing is displayed rather than
taken on trust. `causal_dml` is not in the comparison - it is not a predictive family and the
adapter registry, not a list here, is what decides that.
```python
if EXECUTION_TIER == "canonical":
excluded = require_declared_menu_coverage(analysis, case_study=CASE_STUDY)
else:
excluded = analysis.clear()
excluded
```
```python
analysis.sort("label", "family", "config_name", "checkpoint_value")
```
## Interpretation boundaries
The table compares ranking diagnostics under the declared walk-forward protocol. The backtest
engine supplies portfolio returns, transaction costs, contract sizing, and roll execution before
selection.
The distinction is worth stating as a rule rather than as a caveat: **every quantity in this
notebook is a property of the predictions, and every quantity that decides anything is a
property of a portfolio.** A prediction has no size, no holding period, no financing and no
execution price. Those enter in `13_backtest`, and they are capable of reordering the table
below completely - which is why a leaderboard here is a hypothesis about the backtest rather
than a preview of it.
Conformal weighting, when used later as an allocator, calibrates chronologically from prior
validation observations. It uses the calibration-window scale and the finite-sample higher order
statistic. It does not fit a scale on the evaluation fold or pool all folds before calibration.
## Causal diagnostics
Double machine learning answers a different question from prediction. The treatment effect is
conditioned on the configured confounders, and HAC uncertainty follows the decision-time order.
A covariance-estimator failure is not relabeled as HAC. The shared runner must return a finite HAC
standard error for a result to be complete.
**The `refutation_p` column below is not evidence that these effects survived a placebo test.**
The refutation permutes contiguous blocks within each product, and the shared runner sizes those
blocks as `max(label_buffer, treatment_window)`. `causal.treatment_window` is 1 here, so the label
buffer binds and the registered rows carry a 21-period block for `fwd_ret_21d` and a 5-period
block for `fwd_ret_5d`. Neither length is a property of `carry_pct`, whose own persistence the
cell below measures on this case study's feature panel: the autocorrelation is pooled within
product, each product demeaned before pooling so a level difference between products cannot
stand in for persistence within one, and on one row per product-session, because the block
counts sessions.
**The two blocks sit at very different points on that profile, so the concern bears much more
on one label than the other.** Read the block lengths against the autocorrelation at those
lags rather than against the half-life: the decay is slower than the AR(1) half-life implies -
an AR(1) with this lag-1 value would sit at 0.38 by lag 5 and 0.02 by lag 21, where the panel
is at 0.52 and 0.14 - so the half-life is a lower bound on persistence, not the yardstick for
the block. At the 5-session block used for `fwd_ret_5d` the autocorrelation is still 0.52, so that
block leaves real dependence unpreserved; the placebo is a weaker opponent than the truth and
`fwd_ret_5d`'s empirical p-value is biased toward zero by some amount this notebook does not
quantify. At the 21-session block used for `fwd_ret_21d` it is 0.14, and indistinguishable
from zero by lag 63, so that block spans most of the dependence and the concern is
correspondingly weaker there.
**That is no longer what the column reports, and the reason is worth following.** `fwd_ret_5d`
used to sit at 0.0396 and `fwd_ret_21d` at 0.0099, which is 1/101 and the floor 100 draws can
report. The refutation now compares HAC t-statistics rather than raw effects, because a permuted
treatment is not predictable from the controls, its residual keeps nearly all its variance, and
that variance is the denominator of the second-stage effect - so every placebo effect was divided
by a larger number than the observed one. Correcting that moved `fwd_ret_5d` to 0.5545 and
`fwd_ret_21d` to 0.2673, both `Fails`, on an identical fit. The block-length argument above is a
separate, uncorrected narrowing and it bears mainly on `fwd_ret_5d`; either way it is no longer
visible in these two numbers. Read the DML point estimate and its HAC standard error. The
refutation column is recorded for completeness and carries no evidence here.
```python
# The panel carries one row per contract position, so a product-session appears up to three
# times with the same carry_pct. Lagging without de-duplicating steps ~2.78 rows per session
# and reports a persistence profile stretched by that factor. The block the refutation permutes
# counts sessions - run_dml_analysis requires strictly increasing timestamps within a product -
# so sessions are the scale the two have to be compared on.
carry = (
pl.read_parquet(get_case_study_dir(CASE_STUDY) / "features" / "financial.parquet")
.select(["product", "timestamp", "carry_pct"])
.drop_nulls()
.unique(subset=["product", "timestamp"])
.sort(["product", "timestamp"])
)
autocorr = []
for lag in (1, 5, 21, 63):
paired = (
carry.with_columns(pl.col("carry_pct").shift(lag).over("product").alias("lagged"))
.drop_nulls()
.with_columns(
(pl.col("carry_pct") - pl.col("carry_pct").mean().over("product")).alias("x"),
(pl.col("lagged") - pl.col("lagged").mean().over("product")).alias("y"),
)
)
rho = (paired["x"] * paired["y"]).sum() / (
((paired["x"] ** 2).sum() * (paired["y"] ** 2).sum()) ** 0.5
)
autocorr.append({"lag": lag, "autocorrelation": rho, "n_pairs": paired.height})
carry_persistence = pl.DataFrame(autocorr)
half_life = np.log(0.5) / np.log(carry_persistence["autocorrelation"][0])
print(
f"carry_pct within-product pooled autocorrelation, "
f"AR(1) half-life {half_life:.1f} sessions from lag 1"
)
carry_persistence
```
```python
causal_rows = []
for label in ALL_LABELS:
result = CausalResult.one(study, label=label, execution_tier=EXECUTION_TIER)
if not result.complete:
raise RuntimeError(f"causal result for {label} is incomplete")
causal_rows.append({"label": label, "causal_hash": result.hash, **result.metrics})
causal = pl.DataFrame(causal_rows).sort("label")
```
```python
causal
```
## What proceeds to backtesting
All complete prediction rows proceed. The next notebook passes the selected Polars rows directly
to the shared backtest call, publishes product-keyed decisions, and records contract, roll, price,
and prediction lineage. Validation backtest Sharpe, with the prediction checkpoint included in the
configuration identity, is the selection statistic.Reproduit dans son intégralité avec attribution, conformément à la licence de la source. Licence: MIT
Ce résumé a été rédigé par l’agent de recherche de Stratmill à partir de la source originale ; il n’en est pas une copie.