Passer au contenu
Tous les documents de la bibliothèque

Regroupement de régimes de marché par distance de Wasserstein

Article Machine Learning for Trading

Résumé

Le notebook traite chaque fenêtre de rendements comme une distribution empirique et regroupe les fenêtres selon la distance de Wasserstein unidimensionnelle. Pour des échantillons de même taille, le tri fournit l’appariement optimal des quantiles ; la distance reflète donc les différences entre distributions de rendements, notamment le positionnement des queues, que la moyenne et la variance peuvent ne pas révéler. Les barycentres sont calculés quantile par quantile, avec les médianes pour la distance du premier ordre et les moyennes pour celle du deuxième ordre, ce qui permet une procédure de regroupement de type Lloyd.

Le document compare ces regroupements de distributions aux étiquettes de régime et aux méthodes utilisant un petit nombre de moments. Il indique que le regroupement des distributions complètes a mieux retrouvé la référence que k-means appliqué à quatre moments, tandis qu’un mélange gaussien fondé sur ces moments s’en est approché. Il explique aussi comment transformer les étiquettes de fenêtre en variables sans anticipation : attribuer une étiquette à la clôture de la fenêtre, ajuster les centroïdes sur un bloc initial, puis affecter les fenêtres ultérieures à mesure qu’elles arrivent.

Les résultats dépendent de la série et de la configuration illustrées. Les fenêtres ignorent l’ordre des rendements, leur chevauchement n’ajoute pas d’information indépendante et réattribuer les étiquettes ne transforme pas les années d’entraînement utilisées pour ajuster le modèle en données hors échantillon. Une séparation non supervisée ne suffit pas à établir que les groupes correspondent à des régimes économiquement pertinents.

Idées clés

  • Le tri d’échantillons de rendements de même taille transforme la distance de Wasserstein unidimensionnelle en comparaison quantile par quantile.
  • Le barycentre de Wasserstein est la médiane de chaque quantile pour la distance du premier ordre et sa moyenne pour celle du deuxième ordre.
  • Les fenêtres qui se chevauchent augmentent le nombre d’échantillons, mais partagent des observations et n’ajoutent donc pas d’information indépendante.
  • Attribuer les étiquettes de fenêtre à la clôture évite d’utiliser les valeurs futures d’une fenêtre dans les variables des séances antérieures.
  • Ajustez les centroïdes sur un bloc antérieur avant d’affecter les fenêtres ultérieures afin de créer des variables de régime prospectives.

Étiquettes

Texte intégral
# Chapter 7: Defining the Learning Task


# Chapter 7: Defining the Learning Task

The chapter shows that feature research starts long before modeling. It turns raw but validated market data into stable, protocol-safe inputs by enforcing train-only fitting, disciplined outlier handling, explicit representation choices, and visible missing-data rules. The payoff is not cosmetic cleanliness but comparability, auditability, and protection against leakage that would otherwise make later signal evaluation meaningless.

## Learning Objectives

* Build split-aware preprocessing pipelines that produce stable, auditable inputs for label and feature computation.
* Define execution-consistent labels, including fixed-horizon and event-style constructions, and diagnose overlap, resolution behavior, and implied trading intensity.
* Evaluate feature-label bundles fold by fold using appropriate diagnostics for continuous and discrete targets, including stability, shape, and feasibility.
* Screen candidates for implementation feasibility using turnover, break-even cost, and liquidity or capacity checks.
* Account for search bias by defining searched sets, separating exploration from confirmation, and applying appropriate multiple-testing adjustments to fold-level summaries.
* Use mechanism plausibility checks to distinguish potentially stable signal channels from confounded proxies, timing artifacts, and aggregation effects.

## Sections

### 7.1 Data Preprocessing and Encodings

This section shows that feature research starts long before modeling. It turns raw but validated market data into stable, protocol-safe inputs by enforcing train-only fitting, disciplined outlier handling, explicit representation choices, and visible missing-data rules. The payoff is not cosmetic cleanliness but comparability, auditability, and protection against leakage that would otherwise make later signal evaluation meaningless.

- [`01_data_quality_diagnostics`](01_data_quality_diagnostics.ipynb) _(~9 GB RAM)_ — This notebook provides a standardized diagnostic survey across all ML4T datasets. No cleaning is performed—just assessment.
- [`02_preprocessing_pipeline`](02_preprocessing_pipeline.ipynb) _(~12 GB RAM)_ — This notebook implements hands-on cleaning for datasets that need it, plus demonstrates split-aware preprocessing mechanics. The key teaching artifact is the SplitAwarePreprocessor class that prevents lookahead bias.
- [`10_ml4t_library_ecosystem`](10_ml4t_library_ecosystem.ipynb) — This notebook introduces the ml4t library ecosystem used throughout Chapters 7-12: data loaders (ml4t-data), feature computation (ml4t-engineer), and evaluation tools (ml4t-diagnostic). Uses etfs data.

### 7.2 Label Engineering

This section defines the prediction target with trading realism in mind. It moves from execution-consistent fixed-horizon labels to thresholded, adaptive, and event-style labels, then shows how overlap, resolution time, and implied trading intensity shape what the labels actually mean in practice. Readers should care because a one-bar timing mismatch, a poor threshold choice, or ignored overlap can manufacture a signal that will disappear the moment it meets live execution.

- [`03_label_methods`](03_label_methods.ipynb) — This notebook demonstrates all major labeling methods for ML-based trading strategies, with practical examples on real ETF data. It serves as the canonical reference for choosing and configuring labels across all modeling chapters.
- [`04_maximum_favorable_adverse_excursion`](04_maximum_favorable_adverse_excursion.ipynb) — This notebook provides empirical justification for triple-barrier parameter choices. Rather than picking arbitrary barrier widths, we analyze actual price excursions to determine appropriate thresholds.
- `case_studies/etfs/02_labels`
- `case_studies/crypto_perps_funding/02_labels`
- `case_studies/nasdaq100_microstructure/02_labels`
- `case_studies/sp500_equity_option_analytics/02_labels`
- `case_studies/cme_futures/02_labels`
- `case_studies/sp500_options/02_labels`
- `case_studies/us_equities_panel/02_labels`
- `case_studies/us_firm_characteristics/02_labels`
- `case_studies/fx_pairs/02_labels`

### 7.3 Univariate Feature-Label Evaluation

This section builds the book's first serious triage layer for candidate signals. It asks, in order, whether a feature is correct at decision time, associated with the label, shaped in a way that supports plausible mapping to positions, and economically feasible once turnover, costs, and liquidity are considered. Its importance is that it reframes factor research as an auditable screening process rather than a hunt for a single attractive statistic.

- [`05_signal_evaluation`](05_signal_evaluation.ipynb) — This notebook demonstrates single-factor evaluation using Information Coefficient (IC) analysis, quantile returns, and spread metrics. We answer: "Is this factor predictive in the cross-section, and what horizon does it live on?" Uses etfs data.
- [`06_ic_inference`](06_ic_inference.ipynb) — This notebook addresses statistical inference for IC: how confident should we be that a signal's IC is not just noise? We cover HAC adjustment for autocorrelated IC series and block bootstrap for robust confidence intervals.

### 7.4 Search Accounting and Multiple Testing

This section addresses one of the central failure modes of quantitative research: believing the best candidate after a large search. By defining the searched set, separating exploration from confirmation, and introducing FWER and FDR-style corrections, it makes clear that significance depends on how many variants were tried and how they were chosen. Readers should care because this is where the chapter turns anti-overfitting from a slogan into a concrete research accounting discipline.

- [`07_multiple_testing`](07_multiple_testing.ipynb) — This notebook addresses the factor zoo problem: when testing many signals, even with proper inference, the "best" will be inflated by selection bias. We cover FDR control and complexity-aware corrections.

### 7.5 From Correlation to Causality

This section introduces causal thinking as a falsification layer for features that survived statistical screening. Using small DAGs and mechanism plausibility checks, it asks whether a signal reflects a durable channel or merely a confounded proxy, timing artifact, or aggregation mistake. The value for the reader is not that it "proves causality," but that it helps reject weak stories early and carry better-specified hypotheses into later multivariate work.

- [`08_causal_sanity_checks`](08_causal_sanity_checks.ipynb) — This notebook implements lightweight falsification tests for feature evaluation. These tests complement the correlation-based IC analysis from Section 7.3 by checking whether a feature-outcome association is consistent with a proposed mechanism.

## Running the Notebooks

```bash
# From the repository root
uv run python 07_defining_the_learning_task/<notebook>.py

# Test mode (reduced data via Papermill)
uv run pytest tests/test_chapter_notebooks.py -v -k "07_defining_the_learning_task"
```

### Memory and runtime callouts

> Memory: `01_data_quality_diagnostics` peaks at ~9 GB RSS; recommend ≥16 GB system RAM.
> Memory: `02_preprocessing_pipeline` peaks at ~12 GB RSS; recommend ≥24 GB system RAM.
> Runtime: `08_causal_sanity_checks` takes ~6 minutes (200-permutation null across multiple features and horizons).

## References

- [A Simple Sequentially Rejective Multiple Test Procedure on JSTOR](https://www.jstor.org/stable/4615733?seq=1).
- **Clifford S. Asness et al.** (2013). [Value and Momentum Everywhere](https://www.jstor.org/stable/42002613). *The Journal of Finance*.
- **David H. Bailey and Marcos Lopez de Prado** (2014). [The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality](https://doi.org/10.2139/ssrn.2460551).
- **David H. Bailey et al.** (2015). [The Probability of Backtest Overfitting](https://doi.org/10.2139/ssrn.2326253).
- **Yoav Benjamini and Yosef Hochberg** (1995). [Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing](https://www.jstor.org/stable/2346101). *Journal of the Royal Statistical Society. Series B (Methodological)*.
- **Andrew Y. Chen and Tom Zimmermann** (2021). [Open Source Cross-Sectional Asset Pricing](https://doi.org/10.2139/ssrn.3604626).
- **Carlos Cinelli and Chad Hazlett** (2020). [Making sense of sensitivity: extending omitted variable bias](https://www.jstor.org/stable/26895067). *Journal of the Royal Statistical Society. Series B (Statistical Methodology)*.
- **Kent Daniel and Tobias J. Moskowitz** (2016). [Momentum crashes](https://doi.org/10.1016/j.jfineco.2015.12.002). *Journal of Financial Economics*.
- **Paul Glasserman et al.** (2025). [Does Overnight News Explain Overnight Returns?](https://doi.org/10.48550/arXiv.2507.04481).
- **Richard C.. Grinold and Ronald N.. Kahn** (2000). Active portfolio management: A quantitative approach for providing superior returns and controlling risk. *McGraw-Hill*.
- **Campbell R. Harvey et al.** (2016). [...and the Cross-Section of Expected Returns](https://doi.org/10.1093/rfs/hhv059). *Review of Financial Studies*.
- **Kewei Hou et al.** (2020). [Replicating Anomalies](https://doi.org/10.1093/rfs/hhy131). *The Review of Financial Studies*.
- **Narasimhan Jegadeesh and Sheridan Titman** (1993). [Returns to Buying Winners and Selling Losers: Implications for Stock Market Efficiency](https://doi.org/10.1111/j.1540-6261.1993.tb04702.x). *The Journal of Finance*.
- **Theis Ingerslev Jensen et al.** (2022). Is There a Replication Crisis in Finance?.
- **R. David McLean and Jeffrey Pontiff** (2016). [Does Academic Research Destroy Stock Return Predictability?](https://doi.org/10.1111/jofi.12365). *Journal of Finance*.
- **Whitney K. Newey and Kenneth D. West** (1986). [A Simple, Positive Semi-Definite, Heteroskedasticity and AutocorrelationConsistent Covariance Matrix](https://papers.ssrn.com/abstract=225071).
- **Judea Pearl** (2019). [The seven tools of causal inference, with reflections on machine learning](https://doi.org/10.1145/3241036). *Communications of the ACM*.
- **Marcos Lopez de Prado** (2018). Advances in Financial Machine Learning. *John Wiley & Sons*.

Reproduit dans son intégralité avec attribution, conformément à la licence de la source. Licence: MIT

Ce résumé a été rédigé par l’agent de recherche de Stratmill à partir de la source originale ; il n’en est pas une copie.