Agrupamento de regimes de mercado por distribuições de Wasserstein
Resumo
O notebook trata cada janela de retornos como uma distribuição empírica e agrupa as janelas segundo a distância de Wasserstein unidimensional. Para amostras do mesmo tamanho, a ordenação produz a correspondência ótima entre quantis; assim, a distância reflete diferenças em toda a distribuição dos retornos, incluindo a posição das caudas, que média e variância podem não captar. Seus baricentros são calculados quantil a quantil, usando medianas para a distância de primeira ordem e médias para a versão de segunda ordem, o que permite um procedimento de agrupamento no estilo Lloyd.
O documento compara esses agrupamentos distributivos com rótulos de regimes e métodos que usam um pequeno conjunto de momentos. Relata que agrupar as distribuições completas recuperou melhor a referência do que o k-means aplicado a quatro momentos, enquanto uma mistura gaussiana desses momentos chegou perto. Também discute como converter rótulos de janelas em características sem olhar para o futuro: atribuir um rótulo no fechamento da janela, ajustar centróides em um bloco inicial e classificar as janelas posteriores em sequência.
Os resultados são específicos à série e à configuração ilustradas. As janelas descartam a ordem dos retornos, a sobreposição não acrescenta informação independente e os anos de treinamento usados no ajuste não se tornam fora da amostra ao reatribuir rótulos. A separação não supervisionada, por si só, não comprova que os agrupamentos correspondam a regimes economicamente relevantes.
Ideias principais
- A ordenação de amostras de retornos do mesmo tamanho torna a distância de Wasserstein unidimensional uma comparação quantil a quantil.
- O baricentro de Wasserstein é a mediana por quantil para a distância de primeira ordem e a média para a de segunda ordem.
- Janelas sobrepostas criam mais amostras, mas compartilham observações e, portanto, não acrescentam informação independente.
- Atribuir rótulos às janelas no fechamento evita usar valores futuros da janela em características de sessões anteriores.
- Ajuste os centróides em um bloco anterior antes de atribuir as janelas posteriores para criar características de regimes prospectivas.
Tags
Texto completo
# Chapter 7: Defining the Learning Task # Chapter 7: Defining the Learning Task The chapter shows that feature research starts long before modeling. It turns raw but validated market data into stable, protocol-safe inputs by enforcing train-only fitting, disciplined outlier handling, explicit representation choices, and visible missing-data rules. The payoff is not cosmetic cleanliness but comparability, auditability, and protection against leakage that would otherwise make later signal evaluation meaningless. ## Learning Objectives * Build split-aware preprocessing pipelines that produce stable, auditable inputs for label and feature computation. * Define execution-consistent labels, including fixed-horizon and event-style constructions, and diagnose overlap, resolution behavior, and implied trading intensity. * Evaluate feature-label bundles fold by fold using appropriate diagnostics for continuous and discrete targets, including stability, shape, and feasibility. * Screen candidates for implementation feasibility using turnover, break-even cost, and liquidity or capacity checks. * Account for search bias by defining searched sets, separating exploration from confirmation, and applying appropriate multiple-testing adjustments to fold-level summaries. * Use mechanism plausibility checks to distinguish potentially stable signal channels from confounded proxies, timing artifacts, and aggregation effects. ## Sections ### 7.1 Data Preprocessing and Encodings This section shows that feature research starts long before modeling. It turns raw but validated market data into stable, protocol-safe inputs by enforcing train-only fitting, disciplined outlier handling, explicit representation choices, and visible missing-data rules. The payoff is not cosmetic cleanliness but comparability, auditability, and protection against leakage that would otherwise make later signal evaluation meaningless. - [`01_data_quality_diagnostics`](01_data_quality_diagnostics.ipynb) _(~9 GB RAM)_ — This notebook provides a standardized diagnostic survey across all ML4T datasets. No cleaning is performed—just assessment. - [`02_preprocessing_pipeline`](02_preprocessing_pipeline.ipynb) _(~12 GB RAM)_ — This notebook implements hands-on cleaning for datasets that need it, plus demonstrates split-aware preprocessing mechanics. The key teaching artifact is the SplitAwarePreprocessor class that prevents lookahead bias. - [`10_ml4t_library_ecosystem`](10_ml4t_library_ecosystem.ipynb) — This notebook introduces the ml4t library ecosystem used throughout Chapters 7-12: data loaders (ml4t-data), feature computation (ml4t-engineer), and evaluation tools (ml4t-diagnostic). Uses etfs data. ### 7.2 Label Engineering This section defines the prediction target with trading realism in mind. It moves from execution-consistent fixed-horizon labels to thresholded, adaptive, and event-style labels, then shows how overlap, resolution time, and implied trading intensity shape what the labels actually mean in practice. Readers should care because a one-bar timing mismatch, a poor threshold choice, or ignored overlap can manufacture a signal that will disappear the moment it meets live execution. - [`03_label_methods`](03_label_methods.ipynb) — This notebook demonstrates all major labeling methods for ML-based trading strategies, with practical examples on real ETF data. It serves as the canonical reference for choosing and configuring labels across all modeling chapters. - [`04_maximum_favorable_adverse_excursion`](04_maximum_favorable_adverse_excursion.ipynb) — This notebook provides empirical justification for triple-barrier parameter choices. Rather than picking arbitrary barrier widths, we analyze actual price excursions to determine appropriate thresholds. - `case_studies/etfs/02_labels` - `case_studies/crypto_perps_funding/02_labels` - `case_studies/nasdaq100_microstructure/02_labels` - `case_studies/sp500_equity_option_analytics/02_labels` - `case_studies/cme_futures/02_labels` - `case_studies/sp500_options/02_labels` - `case_studies/us_equities_panel/02_labels` - `case_studies/us_firm_characteristics/02_labels` - `case_studies/fx_pairs/02_labels` ### 7.3 Univariate Feature-Label Evaluation This section builds the book's first serious triage layer for candidate signals. It asks, in order, whether a feature is correct at decision time, associated with the label, shaped in a way that supports plausible mapping to positions, and economically feasible once turnover, costs, and liquidity are considered. Its importance is that it reframes factor research as an auditable screening process rather than a hunt for a single attractive statistic. - [`05_signal_evaluation`](05_signal_evaluation.ipynb) — This notebook demonstrates single-factor evaluation using Information Coefficient (IC) analysis, quantile returns, and spread metrics. We answer: "Is this factor predictive in the cross-section, and what horizon does it live on?" Uses etfs data. - [`06_ic_inference`](06_ic_inference.ipynb) — This notebook addresses statistical inference for IC: how confident should we be that a signal's IC is not just noise? We cover HAC adjustment for autocorrelated IC series and block bootstrap for robust confidence intervals. ### 7.4 Search Accounting and Multiple Testing This section addresses one of the central failure modes of quantitative research: believing the best candidate after a large search. By defining the searched set, separating exploration from confirmation, and introducing FWER and FDR-style corrections, it makes clear that significance depends on how many variants were tried and how they were chosen. Readers should care because this is where the chapter turns anti-overfitting from a slogan into a concrete research accounting discipline. - [`07_multiple_testing`](07_multiple_testing.ipynb) — This notebook addresses the factor zoo problem: when testing many signals, even with proper inference, the "best" will be inflated by selection bias. We cover FDR control and complexity-aware corrections. ### 7.5 From Correlation to Causality This section introduces causal thinking as a falsification layer for features that survived statistical screening. Using small DAGs and mechanism plausibility checks, it asks whether a signal reflects a durable channel or merely a confounded proxy, timing artifact, or aggregation mistake. The value for the reader is not that it "proves causality," but that it helps reject weak stories early and carry better-specified hypotheses into later multivariate work. - [`08_causal_sanity_checks`](08_causal_sanity_checks.ipynb) — This notebook implements lightweight falsification tests for feature evaluation. These tests complement the correlation-based IC analysis from Section 7.3 by checking whether a feature-outcome association is consistent with a proposed mechanism. ## Running the Notebooks ```bash # From the repository root uv run python 07_defining_the_learning_task/<notebook>.py # Test mode (reduced data via Papermill) uv run pytest tests/test_chapter_notebooks.py -v -k "07_defining_the_learning_task" ``` ### Memory and runtime callouts > Memory: `01_data_quality_diagnostics` peaks at ~9 GB RSS; recommend ≥16 GB system RAM. > Memory: `02_preprocessing_pipeline` peaks at ~12 GB RSS; recommend ≥24 GB system RAM. > Runtime: `08_causal_sanity_checks` takes ~6 minutes (200-permutation null across multiple features and horizons). ## References - [A Simple Sequentially Rejective Multiple Test Procedure on JSTOR](https://www.jstor.org/stable/4615733?seq=1). - **Clifford S. Asness et al.** (2013). [Value and Momentum Everywhere](https://www.jstor.org/stable/42002613). *The Journal of Finance*. - **David H. Bailey and Marcos Lopez de Prado** (2014). [The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality](https://doi.org/10.2139/ssrn.2460551). - **David H. Bailey et al.** (2015). [The Probability of Backtest Overfitting](https://doi.org/10.2139/ssrn.2326253). - **Yoav Benjamini and Yosef Hochberg** (1995). [Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing](https://www.jstor.org/stable/2346101). *Journal of the Royal Statistical Society. Series B (Methodological)*. - **Andrew Y. Chen and Tom Zimmermann** (2021). [Open Source Cross-Sectional Asset Pricing](https://doi.org/10.2139/ssrn.3604626). - **Carlos Cinelli and Chad Hazlett** (2020). [Making sense of sensitivity: extending omitted variable bias](https://www.jstor.org/stable/26895067). *Journal of the Royal Statistical Society. Series B (Statistical Methodology)*. - **Kent Daniel and Tobias J. Moskowitz** (2016). [Momentum crashes](https://doi.org/10.1016/j.jfineco.2015.12.002). *Journal of Financial Economics*. - **Paul Glasserman et al.** (2025). [Does Overnight News Explain Overnight Returns?](https://doi.org/10.48550/arXiv.2507.04481). - **Richard C.. Grinold and Ronald N.. Kahn** (2000). Active portfolio management: A quantitative approach for providing superior returns and controlling risk. *McGraw-Hill*. - **Campbell R. Harvey et al.** (2016). [...and the Cross-Section of Expected Returns](https://doi.org/10.1093/rfs/hhv059). *Review of Financial Studies*. - **Kewei Hou et al.** (2020). [Replicating Anomalies](https://doi.org/10.1093/rfs/hhy131). *The Review of Financial Studies*. - **Narasimhan Jegadeesh and Sheridan Titman** (1993). [Returns to Buying Winners and Selling Losers: Implications for Stock Market Efficiency](https://doi.org/10.1111/j.1540-6261.1993.tb04702.x). *The Journal of Finance*. - **Theis Ingerslev Jensen et al.** (2022). Is There a Replication Crisis in Finance?. - **R. David McLean and Jeffrey Pontiff** (2016). [Does Academic Research Destroy Stock Return Predictability?](https://doi.org/10.1111/jofi.12365). *Journal of Finance*. - **Whitney K. Newey and Kenneth D. West** (1986). [A Simple, Positive Semi-Definite, Heteroskedasticity and AutocorrelationConsistent Covariance Matrix](https://papers.ssrn.com/abstract=225071). - **Judea Pearl** (2019). [The seven tools of causal inference, with reflections on machine learning](https://doi.org/10.1145/3241036). *Communications of the ACM*. - **Marcos Lopez de Prado** (2018). Advances in Financial Machine Learning. *John Wiley & Sons*.
Exibido na íntegra, com atribuição conforme a licença da fonte. Licença: MIT
Este resumo foi escrito pelo agente de pesquisa da Stratmill com base no original; não é uma cópia da fonte.