Wasserstein-Clustering für verteilungsbasierte Marktregime
Zusammenfassung
Das Notebook behandelt jedes Renditefenster als empirische Verteilung und clustert Fenster anhand der eindimensionalen Wasserstein-Distanz. Bei gleich großen Stichproben ergibt das Sortieren die optimale Zuordnung der Quantile. Die Distanz bildet daher Unterschiede in der gesamten Renditeverteilung ab, einschließlich der Lage der Verteilungsschwänze, die Mittelwert und Varianz übersehen können. Die Baryzentren werden Quantil für Quantil berechnet: Für die Distanz erster Ordnung werden Mediane, für die Version zweiter Ordnung Mittelwerte verwendet. Das ermöglicht ein Lloyd-ähnliches Clustering-Verfahren.
Das Dokument vergleicht diese Verteilungscluster mit Regime-Labels und Verfahren, die eine kleine Zahl von Momenten verwenden. Es berichtet, dass das Clustering der vollständigen Verteilungen die Referenz besser wiederherstellte als k-means mit vier Momenten, während eine Gaußsche Mischung auf Basis dieser Momente nahe herankam. Außerdem wird beschrieben, wie Fenster-Labels ohne Blick in die Zukunft in Merkmale umgewandelt werden: Das Label wird am Fensterschluss festgehalten, die Zentroiden werden auf einem anfänglichen Block angepasst und spätere Fenster anschließend vorwärts zugeordnet.
Die Ergebnisse gelten für die dargestellte Reihe und Konfiguration. Fenster verwerfen die Reihenfolge der Renditen; durch Überlappung entsteht keine unabhängige Information, und Trainingsjahre werden nicht dadurch zu Out-of-Sample-Daten, dass Labels neu festgehalten werden. Eine Trennung durch unüberwachtes Lernen allein belegt nicht, dass Cluster wirtschaftlich sinnvollen Regimen entsprechen.
Kernaussagen
- Bei gleich großen Renditestichproben macht Sortieren die eindimensionale Wasserstein-Distanz zu einem Quantilvergleich.
- Das Wasserstein-Baryzentrum ist bei der Distanz erster Ordnung der Median je Quantil und bei der Distanz zweiter Ordnung der Mittelwert je Quantil.
- Überlappende Fenster erhöhen die Stichprobenzahl, teilen aber Beobachtungen und schaffen daher keine unabhängige Information.
- Werden Fenster-Labels am Fensterschluss zugewiesen, gelangen keine späteren Fensterwerte in Merkmale früherer Sitzungen.
- Passen Sie Zentroiden auf einem früheren Block an, bevor Sie spätere Fenster zuordnen, um vorwärtsgerichtete Regimemerkmale zu erstellen.
Schlagwörter
Volltext
# Chapter 7: Defining the Learning Task # Chapter 7: Defining the Learning Task The chapter shows that feature research starts long before modeling. It turns raw but validated market data into stable, protocol-safe inputs by enforcing train-only fitting, disciplined outlier handling, explicit representation choices, and visible missing-data rules. The payoff is not cosmetic cleanliness but comparability, auditability, and protection against leakage that would otherwise make later signal evaluation meaningless. ## Learning Objectives * Build split-aware preprocessing pipelines that produce stable, auditable inputs for label and feature computation. * Define execution-consistent labels, including fixed-horizon and event-style constructions, and diagnose overlap, resolution behavior, and implied trading intensity. * Evaluate feature-label bundles fold by fold using appropriate diagnostics for continuous and discrete targets, including stability, shape, and feasibility. * Screen candidates for implementation feasibility using turnover, break-even cost, and liquidity or capacity checks. * Account for search bias by defining searched sets, separating exploration from confirmation, and applying appropriate multiple-testing adjustments to fold-level summaries. * Use mechanism plausibility checks to distinguish potentially stable signal channels from confounded proxies, timing artifacts, and aggregation effects. ## Sections ### 7.1 Data Preprocessing and Encodings This section shows that feature research starts long before modeling. It turns raw but validated market data into stable, protocol-safe inputs by enforcing train-only fitting, disciplined outlier handling, explicit representation choices, and visible missing-data rules. The payoff is not cosmetic cleanliness but comparability, auditability, and protection against leakage that would otherwise make later signal evaluation meaningless. - [`01_data_quality_diagnostics`](01_data_quality_diagnostics.ipynb) _(~9 GB RAM)_ — This notebook provides a standardized diagnostic survey across all ML4T datasets. No cleaning is performed—just assessment. - [`02_preprocessing_pipeline`](02_preprocessing_pipeline.ipynb) _(~12 GB RAM)_ — This notebook implements hands-on cleaning for datasets that need it, plus demonstrates split-aware preprocessing mechanics. The key teaching artifact is the SplitAwarePreprocessor class that prevents lookahead bias. - [`10_ml4t_library_ecosystem`](10_ml4t_library_ecosystem.ipynb) — This notebook introduces the ml4t library ecosystem used throughout Chapters 7-12: data loaders (ml4t-data), feature computation (ml4t-engineer), and evaluation tools (ml4t-diagnostic). Uses etfs data. ### 7.2 Label Engineering This section defines the prediction target with trading realism in mind. It moves from execution-consistent fixed-horizon labels to thresholded, adaptive, and event-style labels, then shows how overlap, resolution time, and implied trading intensity shape what the labels actually mean in practice. Readers should care because a one-bar timing mismatch, a poor threshold choice, or ignored overlap can manufacture a signal that will disappear the moment it meets live execution. - [`03_label_methods`](03_label_methods.ipynb) — This notebook demonstrates all major labeling methods for ML-based trading strategies, with practical examples on real ETF data. It serves as the canonical reference for choosing and configuring labels across all modeling chapters. - [`04_maximum_favorable_adverse_excursion`](04_maximum_favorable_adverse_excursion.ipynb) — This notebook provides empirical justification for triple-barrier parameter choices. Rather than picking arbitrary barrier widths, we analyze actual price excursions to determine appropriate thresholds. - `case_studies/etfs/02_labels` - `case_studies/crypto_perps_funding/02_labels` - `case_studies/nasdaq100_microstructure/02_labels` - `case_studies/sp500_equity_option_analytics/02_labels` - `case_studies/cme_futures/02_labels` - `case_studies/sp500_options/02_labels` - `case_studies/us_equities_panel/02_labels` - `case_studies/us_firm_characteristics/02_labels` - `case_studies/fx_pairs/02_labels` ### 7.3 Univariate Feature-Label Evaluation This section builds the book's first serious triage layer for candidate signals. It asks, in order, whether a feature is correct at decision time, associated with the label, shaped in a way that supports plausible mapping to positions, and economically feasible once turnover, costs, and liquidity are considered. Its importance is that it reframes factor research as an auditable screening process rather than a hunt for a single attractive statistic. - [`05_signal_evaluation`](05_signal_evaluation.ipynb) — This notebook demonstrates single-factor evaluation using Information Coefficient (IC) analysis, quantile returns, and spread metrics. We answer: "Is this factor predictive in the cross-section, and what horizon does it live on?" Uses etfs data. - [`06_ic_inference`](06_ic_inference.ipynb) — This notebook addresses statistical inference for IC: how confident should we be that a signal's IC is not just noise? We cover HAC adjustment for autocorrelated IC series and block bootstrap for robust confidence intervals. ### 7.4 Search Accounting and Multiple Testing This section addresses one of the central failure modes of quantitative research: believing the best candidate after a large search. By defining the searched set, separating exploration from confirmation, and introducing FWER and FDR-style corrections, it makes clear that significance depends on how many variants were tried and how they were chosen. Readers should care because this is where the chapter turns anti-overfitting from a slogan into a concrete research accounting discipline. - [`07_multiple_testing`](07_multiple_testing.ipynb) — This notebook addresses the factor zoo problem: when testing many signals, even with proper inference, the "best" will be inflated by selection bias. We cover FDR control and complexity-aware corrections. ### 7.5 From Correlation to Causality This section introduces causal thinking as a falsification layer for features that survived statistical screening. Using small DAGs and mechanism plausibility checks, it asks whether a signal reflects a durable channel or merely a confounded proxy, timing artifact, or aggregation mistake. The value for the reader is not that it "proves causality," but that it helps reject weak stories early and carry better-specified hypotheses into later multivariate work. - [`08_causal_sanity_checks`](08_causal_sanity_checks.ipynb) — This notebook implements lightweight falsification tests for feature evaluation. These tests complement the correlation-based IC analysis from Section 7.3 by checking whether a feature-outcome association is consistent with a proposed mechanism. ## Running the Notebooks ```bash # From the repository root uv run python 07_defining_the_learning_task/<notebook>.py # Test mode (reduced data via Papermill) uv run pytest tests/test_chapter_notebooks.py -v -k "07_defining_the_learning_task" ``` ### Memory and runtime callouts > Memory: `01_data_quality_diagnostics` peaks at ~9 GB RSS; recommend ≥16 GB system RAM. > Memory: `02_preprocessing_pipeline` peaks at ~12 GB RSS; recommend ≥24 GB system RAM. > Runtime: `08_causal_sanity_checks` takes ~6 minutes (200-permutation null across multiple features and horizons). ## References - [A Simple Sequentially Rejective Multiple Test Procedure on JSTOR](https://www.jstor.org/stable/4615733?seq=1). - **Clifford S. Asness et al.** (2013). [Value and Momentum Everywhere](https://www.jstor.org/stable/42002613). *The Journal of Finance*. - **David H. Bailey and Marcos Lopez de Prado** (2014). [The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality](https://doi.org/10.2139/ssrn.2460551). - **David H. Bailey et al.** (2015). [The Probability of Backtest Overfitting](https://doi.org/10.2139/ssrn.2326253). - **Yoav Benjamini and Yosef Hochberg** (1995). [Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing](https://www.jstor.org/stable/2346101). *Journal of the Royal Statistical Society. Series B (Methodological)*. - **Andrew Y. Chen and Tom Zimmermann** (2021). [Open Source Cross-Sectional Asset Pricing](https://doi.org/10.2139/ssrn.3604626). - **Carlos Cinelli and Chad Hazlett** (2020). [Making sense of sensitivity: extending omitted variable bias](https://www.jstor.org/stable/26895067). *Journal of the Royal Statistical Society. Series B (Statistical Methodology)*. - **Kent Daniel and Tobias J. Moskowitz** (2016). [Momentum crashes](https://doi.org/10.1016/j.jfineco.2015.12.002). *Journal of Financial Economics*. - **Paul Glasserman et al.** (2025). [Does Overnight News Explain Overnight Returns?](https://doi.org/10.48550/arXiv.2507.04481). - **Richard C.. Grinold and Ronald N.. Kahn** (2000). Active portfolio management: A quantitative approach for providing superior returns and controlling risk. *McGraw-Hill*. - **Campbell R. Harvey et al.** (2016). [...and the Cross-Section of Expected Returns](https://doi.org/10.1093/rfs/hhv059). *Review of Financial Studies*. - **Kewei Hou et al.** (2020). [Replicating Anomalies](https://doi.org/10.1093/rfs/hhy131). *The Review of Financial Studies*. - **Narasimhan Jegadeesh and Sheridan Titman** (1993). [Returns to Buying Winners and Selling Losers: Implications for Stock Market Efficiency](https://doi.org/10.1111/j.1540-6261.1993.tb04702.x). *The Journal of Finance*. - **Theis Ingerslev Jensen et al.** (2022). Is There a Replication Crisis in Finance?. - **R. David McLean and Jeffrey Pontiff** (2016). [Does Academic Research Destroy Stock Return Predictability?](https://doi.org/10.1111/jofi.12365). *Journal of Finance*. - **Whitney K. Newey and Kenneth D. West** (1986). [A Simple, Positive Semi-Definite, Heteroskedasticity and AutocorrelationConsistent Covariance Matrix](https://papers.ssrn.com/abstract=225071). - **Judea Pearl** (2019). [The seven tools of causal inference, with reflections on machine learning](https://doi.org/10.1145/3241036). *Communications of the ACM*. - **Marcos Lopez de Prado** (2018). Advances in Financial Machine Learning. *John Wiley & Sons*.
Vollständig mit Quellenangabe unter der Lizenz der Quelle angezeigt. Lizenz: MIT
Diese Zusammenfassung wurde vom Research-Agenten von Stratmill anhand des Originals verfasst; sie ist keine Kopie der Quelle.