Перейти к содержимому
Все документы библиотеки

Кластеризация рыночных режимов по расстоянию Васерштейна

Статья Machine Learning for Trading

Сводка

В тетради каждое окно доходности рассматривается как эмпирическое распределение, а окна кластеризуются по одномерному расстоянию Васерштейна. Для выборок одинакового размера сортировка даёт оптимальное сопоставление квантилей; поэтому расстояние отражает различия во всём распределении доходности, включая положение хвостов, которые могут не улавливаться средним и дисперсией. Барицентры вычисляются по каждому квантилю: используются медианы для расстояния первого порядка и средние для варианта второго порядка, что позволяет применить процедуру кластеризации в стиле Ллойда.

Документ сравнивает кластеры распределений с метками режимов и методами на небольшом наборе моментов. Сообщается, что кластеризация полных распределений лучше восстановила эталон, чем k-means по четырём моментам, тогда как смесь Гаусса на этих моментах показала близкий результат. Также обсуждается преобразование меток окон в признаки без заглядывания в будущее: присваивать метку при закрытии окна, подбирать центроиды на начальном блоке и затем относить к ним более поздние окна.

Результаты относятся только к показанному ряду и настройке. Окна отбрасывают порядок доходностей, перекрытие не добавляет независимой информации, а обучающие годы нельзя превратить во вневыборочные повторной расстановкой меток. Одно лишь разделение на группы без учителя не доказывает, что кластеры соответствуют экономически осмысленным режимам.

Ключевые идеи

  • Сортировка выборок доходности одинакового размера превращает одномерное расстояние Васерштейна в сравнение квантилей.
  • Барицентр Васерштейна — это поквантильная медиана для расстояния первого порядка и среднее для расстояния второго порядка.
  • Перекрывающиеся окна увеличивают число выборок, но используют общие наблюдения и поэтому не добавляют независимой информации.
  • Присвоение меток окнам при закрытии позволяет не использовать будущие значения окна в признаках для более ранних сессий.
  • Подбери центроиды на более раннем блоке, прежде чем относить к ним последующие окна, чтобы получить признаки режимов с прогнозом вперёд.

Теги

Полный текст
# Chapter 7: Defining the Learning Task


# Chapter 7: Defining the Learning Task

The chapter shows that feature research starts long before modeling. It turns raw but validated market data into stable, protocol-safe inputs by enforcing train-only fitting, disciplined outlier handling, explicit representation choices, and visible missing-data rules. The payoff is not cosmetic cleanliness but comparability, auditability, and protection against leakage that would otherwise make later signal evaluation meaningless.

## Learning Objectives

* Build split-aware preprocessing pipelines that produce stable, auditable inputs for label and feature computation.
* Define execution-consistent labels, including fixed-horizon and event-style constructions, and diagnose overlap, resolution behavior, and implied trading intensity.
* Evaluate feature-label bundles fold by fold using appropriate diagnostics for continuous and discrete targets, including stability, shape, and feasibility.
* Screen candidates for implementation feasibility using turnover, break-even cost, and liquidity or capacity checks.
* Account for search bias by defining searched sets, separating exploration from confirmation, and applying appropriate multiple-testing adjustments to fold-level summaries.
* Use mechanism plausibility checks to distinguish potentially stable signal channels from confounded proxies, timing artifacts, and aggregation effects.

## Sections

### 7.1 Data Preprocessing and Encodings

This section shows that feature research starts long before modeling. It turns raw but validated market data into stable, protocol-safe inputs by enforcing train-only fitting, disciplined outlier handling, explicit representation choices, and visible missing-data rules. The payoff is not cosmetic cleanliness but comparability, auditability, and protection against leakage that would otherwise make later signal evaluation meaningless.

- [`01_data_quality_diagnostics`](01_data_quality_diagnostics.ipynb) _(~9 GB RAM)_ — This notebook provides a standardized diagnostic survey across all ML4T datasets. No cleaning is performed—just assessment.
- [`02_preprocessing_pipeline`](02_preprocessing_pipeline.ipynb) _(~12 GB RAM)_ — This notebook implements hands-on cleaning for datasets that need it, plus demonstrates split-aware preprocessing mechanics. The key teaching artifact is the SplitAwarePreprocessor class that prevents lookahead bias.
- [`10_ml4t_library_ecosystem`](10_ml4t_library_ecosystem.ipynb) — This notebook introduces the ml4t library ecosystem used throughout Chapters 7-12: data loaders (ml4t-data), feature computation (ml4t-engineer), and evaluation tools (ml4t-diagnostic). Uses etfs data.

### 7.2 Label Engineering

This section defines the prediction target with trading realism in mind. It moves from execution-consistent fixed-horizon labels to thresholded, adaptive, and event-style labels, then shows how overlap, resolution time, and implied trading intensity shape what the labels actually mean in practice. Readers should care because a one-bar timing mismatch, a poor threshold choice, or ignored overlap can manufacture a signal that will disappear the moment it meets live execution.

- [`03_label_methods`](03_label_methods.ipynb) — This notebook demonstrates all major labeling methods for ML-based trading strategies, with practical examples on real ETF data. It serves as the canonical reference for choosing and configuring labels across all modeling chapters.
- [`04_maximum_favorable_adverse_excursion`](04_maximum_favorable_adverse_excursion.ipynb) — This notebook provides empirical justification for triple-barrier parameter choices. Rather than picking arbitrary barrier widths, we analyze actual price excursions to determine appropriate thresholds.
- `case_studies/etfs/02_labels`
- `case_studies/crypto_perps_funding/02_labels`
- `case_studies/nasdaq100_microstructure/02_labels`
- `case_studies/sp500_equity_option_analytics/02_labels`
- `case_studies/cme_futures/02_labels`
- `case_studies/sp500_options/02_labels`
- `case_studies/us_equities_panel/02_labels`
- `case_studies/us_firm_characteristics/02_labels`
- `case_studies/fx_pairs/02_labels`

### 7.3 Univariate Feature-Label Evaluation

This section builds the book's first serious triage layer for candidate signals. It asks, in order, whether a feature is correct at decision time, associated with the label, shaped in a way that supports plausible mapping to positions, and economically feasible once turnover, costs, and liquidity are considered. Its importance is that it reframes factor research as an auditable screening process rather than a hunt for a single attractive statistic.

- [`05_signal_evaluation`](05_signal_evaluation.ipynb) — This notebook demonstrates single-factor evaluation using Information Coefficient (IC) analysis, quantile returns, and spread metrics. We answer: "Is this factor predictive in the cross-section, and what horizon does it live on?" Uses etfs data.
- [`06_ic_inference`](06_ic_inference.ipynb) — This notebook addresses statistical inference for IC: how confident should we be that a signal's IC is not just noise? We cover HAC adjustment for autocorrelated IC series and block bootstrap for robust confidence intervals.

### 7.4 Search Accounting and Multiple Testing

This section addresses one of the central failure modes of quantitative research: believing the best candidate after a large search. By defining the searched set, separating exploration from confirmation, and introducing FWER and FDR-style corrections, it makes clear that significance depends on how many variants were tried and how they were chosen. Readers should care because this is where the chapter turns anti-overfitting from a slogan into a concrete research accounting discipline.

- [`07_multiple_testing`](07_multiple_testing.ipynb) — This notebook addresses the factor zoo problem: when testing many signals, even with proper inference, the "best" will be inflated by selection bias. We cover FDR control and complexity-aware corrections.

### 7.5 From Correlation to Causality

This section introduces causal thinking as a falsification layer for features that survived statistical screening. Using small DAGs and mechanism plausibility checks, it asks whether a signal reflects a durable channel or merely a confounded proxy, timing artifact, or aggregation mistake. The value for the reader is not that it "proves causality," but that it helps reject weak stories early and carry better-specified hypotheses into later multivariate work.

- [`08_causal_sanity_checks`](08_causal_sanity_checks.ipynb) — This notebook implements lightweight falsification tests for feature evaluation. These tests complement the correlation-based IC analysis from Section 7.3 by checking whether a feature-outcome association is consistent with a proposed mechanism.

## Running the Notebooks

```bash
# From the repository root
uv run python 07_defining_the_learning_task/<notebook>.py

# Test mode (reduced data via Papermill)
uv run pytest tests/test_chapter_notebooks.py -v -k "07_defining_the_learning_task"
```

### Memory and runtime callouts

> Memory: `01_data_quality_diagnostics` peaks at ~9 GB RSS; recommend ≥16 GB system RAM.
> Memory: `02_preprocessing_pipeline` peaks at ~12 GB RSS; recommend ≥24 GB system RAM.
> Runtime: `08_causal_sanity_checks` takes ~6 minutes (200-permutation null across multiple features and horizons).

## References

- [A Simple Sequentially Rejective Multiple Test Procedure on JSTOR](https://www.jstor.org/stable/4615733?seq=1).
- **Clifford S. Asness et al.** (2013). [Value and Momentum Everywhere](https://www.jstor.org/stable/42002613). *The Journal of Finance*.
- **David H. Bailey and Marcos Lopez de Prado** (2014). [The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality](https://doi.org/10.2139/ssrn.2460551).
- **David H. Bailey et al.** (2015). [The Probability of Backtest Overfitting](https://doi.org/10.2139/ssrn.2326253).
- **Yoav Benjamini and Yosef Hochberg** (1995). [Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing](https://www.jstor.org/stable/2346101). *Journal of the Royal Statistical Society. Series B (Methodological)*.
- **Andrew Y. Chen and Tom Zimmermann** (2021). [Open Source Cross-Sectional Asset Pricing](https://doi.org/10.2139/ssrn.3604626).
- **Carlos Cinelli and Chad Hazlett** (2020). [Making sense of sensitivity: extending omitted variable bias](https://www.jstor.org/stable/26895067). *Journal of the Royal Statistical Society. Series B (Statistical Methodology)*.
- **Kent Daniel and Tobias J. Moskowitz** (2016). [Momentum crashes](https://doi.org/10.1016/j.jfineco.2015.12.002). *Journal of Financial Economics*.
- **Paul Glasserman et al.** (2025). [Does Overnight News Explain Overnight Returns?](https://doi.org/10.48550/arXiv.2507.04481).
- **Richard C.. Grinold and Ronald N.. Kahn** (2000). Active portfolio management: A quantitative approach for providing superior returns and controlling risk. *McGraw-Hill*.
- **Campbell R. Harvey et al.** (2016). [...and the Cross-Section of Expected Returns](https://doi.org/10.1093/rfs/hhv059). *Review of Financial Studies*.
- **Kewei Hou et al.** (2020). [Replicating Anomalies](https://doi.org/10.1093/rfs/hhy131). *The Review of Financial Studies*.
- **Narasimhan Jegadeesh and Sheridan Titman** (1993). [Returns to Buying Winners and Selling Losers: Implications for Stock Market Efficiency](https://doi.org/10.1111/j.1540-6261.1993.tb04702.x). *The Journal of Finance*.
- **Theis Ingerslev Jensen et al.** (2022). Is There a Replication Crisis in Finance?.
- **R. David McLean and Jeffrey Pontiff** (2016). [Does Academic Research Destroy Stock Return Predictability?](https://doi.org/10.1111/jofi.12365). *Journal of Finance*.
- **Whitney K. Newey and Kenneth D. West** (1986). [A Simple, Positive Semi-Definite, Heteroskedasticity and AutocorrelationConsistent Covariance Matrix](https://papers.ssrn.com/abstract=225071).
- **Judea Pearl** (2019). [The seven tools of causal inference, with reflections on machine learning](https://doi.org/10.1145/3241036). *Communications of the ACM*.
- **Marcos Lopez de Prado** (2018). Advances in Financial Machine Learning. *John Wiley & Sons*.

Полный текст с указанием источника опубликован на условиях его лицензии. Лицензия: MIT

Это краткое изложение подготовлено исследовательским агентом Stratmill по оригиналу и не является его копией.