Univers thématique de ETF pour la recherche sur le momentum
Résumé
Cette note sur les données décrit une collection diversifiée de fonds négociés en bourse utilisée dans une stratégie de momentum et dans une série plus large d’exemples de recherche financière. Elle fournit des observations quotidiennes d’ouverture, de plus haut, de plus bas, de clôture et de volume à partir de 2006, regroupées entre styles et secteurs d’actions US, marchés internationaux et émergents, titres à revenu fixe, matières premières, thèmes spécialisés et devises. Le document explique également la source et l’organisation du jeu de données, ses métadonnées et les options de chargement permettant de sélectionner des symboles et des périodes.
La diversité des catégories d’actifs rend cette collection adaptée à la recherche transversale et multi-actifs. Ses utilisations en aval déclarées comprennent le développement de stratégies, l’analyse de caractéristiques, les études de robustesse, l’exécution et la présentation des performances. La note ne présente pas les rendements d’une règle de trading et n’établit pas que les données permettent de construire une stratégie rentable. Les données proviennent de Yahoo Finance, selon des conditions d’utilisation personnelles et pédagogiques déclarées ; la redistribution des données brutes est restreinte, et il convient de tenir compte de la couverture du fournisseur, des ajustements et de la qualité des données pour tirer des conclusions empiriques.
Idées clés
- La collection contient des historiques quotidiens OHLCV pour des FNB 100 répartis entre neuf groupes thématiques.
- La couverture comprend les actions, les titres à revenu fixe, les matières premières, les devises et les marchés internationaux.
- Un schéma commun et les métadonnées des symboles permettent de charger des sous-ensembles par ticker et période.
- Le jeu de données sert à la recherche sur le momentum ainsi qu’à plusieurs exemples d’analyse et d’exécution.
- Les conditions déclarées par la source restreignent la redistribution des observations brutes.
Étiquettes
Texte intégral
# Chapter 5: Synthetic Financial Data # Chapter 5: Synthetic Financial Data Financial research has one realized price path and a research process that searches it repeatedly, which is enough on its own to make a backtest fragile. This chapter asks what it would take to generate more paths worth trusting: what a generator has to preserve, how to tell whether it did, and what the answer costs in privacy and in effort. It works through classical simulation baselines, GANs, diffusion models and LLM-based tabular generation, and it holds each of them to the same evaluation, because the validation protocol matters more here than the choice of architecture. ## Learning Objectives * Explain why trading research is path-limited and how adaptive search and multiple testing can inflate apparent backtest performance. * Use classical simulation baselines, including bootstrap and stochastic volatility models, as interpretable benchmarks for synthetic data generation. * Select a synthetic-data approach that matches the data structure and downstream objective, including learned generators for time series and tabular financial data. * Diagnose generated data using stylized-fact, dependence, and task-based evaluation methods, including Train-Synthetic-Test-Real comparisons. * Assess privacy and generator-specific risks, including leakage, bias amplification, overfitting to the generator, and limited scenario novelty. ## Sections ### 5.1 The Quant's Dilemma This section explains why backtests are fragile even before any generative model enters the picture. Because strategy research is adaptive and path-limited, strong in-sample results may simply reflect favorable history and repeated search. The section then motivates synthetic data as a way to expand robustness analysis beyond one realized market path, while anchoring the discussion in the stylized facts that any useful generator must preserve. ### 5.2 Evaluating Synthetic Financial Data This section sets the standard the rest of the chapter is measured against, which is why it comes before any generator. Synthetic data can look plausible and still be unusable, because the structure that matters in finance sits in rare events, in shifts in dependence, and in conditional dynamics rather than at the centre of the distribution. It introduces the three criteria the chapter uses throughout - fidelity, utility and privacy - and the fact that they trade off against one another rather than improving together. ### 5.3 Classical Simulation Baselines This section gives the chapter intellectual discipline. Rather than jumping straight to deep generators, it shows that bootstrap methods, GBM, jump-diffusion, mean reversion, Heston, and GARCH remain important baselines because they are interpretable, sample-efficient, and easier to validate. Readers should care because these methods define the benchmark a learned generator ought to beat on the diagnostics that actually matter. - [`00_classical_simulation`](00_classical_simulation.ipynb) — Simulates return paths from GBM, jump diffusion, mean reversion, Heston and GARCH, and resamples real returns with the IID, block and stationary bootstraps, then measures which of them reproduce fat tails and volatility clustering. These are the baselines a learned generator has to beat. Uses ETF data. ### 5.4 Generative Model Taxonomy This section provides the conceptual map for the rest of the chapter. It distinguishes discriminative from generative modeling and positions VAEs, GANs, diffusion models, and LLM-based tabular generators as alternative ways to learn joint distributions rather than hand-specify them. Its value is orientation: readers can see early that architecture choice is really about preserving the structure their downstream use case needs. ### 5.5 GANs for Financial Time Series This is the chapter's most differentiated model survey. It moves from the basic adversarial setup to finance-specific variants such as TimeGAN, Tail-GAN, Sig-CWGAN, and GT-GAN, showing how each responds to a concrete weakness of vanilla GANs: temporal structure, tail risk, path fidelity, or irregular timestamps. The section matters because it teaches readers not to ask which GAN is best in general, but which inductive bias matches the task and its failure modes. - [`01_timegan`](01_timegan.ipynb) — This notebook implements TimeGAN (Yoon, Jarrett & van der Schaar, NeurIPS 2019), the foundational architecture for synthetic financial time series generation. Uses multi_stock_data, state_dict, us_equities data. - [`02_tailgan_tail_risk`](02_tailgan_tail_risk.ipynb) — This notebook implements Tail-GAN (Cont, Xu, and Zhang 2022), a GAN architecture that uses differentiable sorting to preserve tail risk characteristics (VaR, ES) in synthetic financial scenarios. Uses etf_returns, etfs, state_dict data. - [`03_sigcwgan_signatures`](03_sigcwgan_signatures.ipynb) — Builds the unconditional Sig-Wasserstein GAN, which replaces the learned discriminator with an analytic distance between expected path signatures, and explains how the conditional Sig-CWGAN extends it. Uses S&P 500 log returns. Needs the `py312` profile; see the callout at the end of this file. - [`04_gtgan_irregular`](04_gtgan_irregular.ipynb) — This notebook implements a GT-GAN-inspired model (based on Jeon et al., NeurIPS 2022) using Neural ODEs to handle time series with naturally irregular timestamps. Uses Chapter 3 NVDA dollar bars (Databento bar sampling). ### 5.6 Diffusion Models for Financial Time Series This section presents diffusion models as a strong, often more stable alternative to adversarial training. It explains the denoising framework, why diffusion can fit financial return structure, how conditional guidance supports regime-aware stress testing, and why Diffusion-TS is a useful reference design for sequential financial data. Readers should care because this is the chapter's clearest candidate for a general-purpose learned generator, but one that still demands careful validation and use-case alignment. - [`05_diffusion_ts`](05_diffusion_ts.ipynb) — This notebook implements Diffusion-TS (Yuan & Qiao, ICLR 2024), a diffusion model that decomposes the denoising prediction into trend (polynomial regression) and seasonal (Fourier basis) components. This interpretable structure encourages the model to separate slow drift from periodic patterns, analogous to classical STL decomposition but learned end-to-end within the diffusion framework. ### 5.7 LLMs for Structured Financial Data This section extends the synthetic-data discussion beyond return series to mixed-type financial tables. By introducing serialization, the GReaT workflow, and constraint-based postprocessing, it shows where LLMs can be practical for credit, customer, and fundamental data. The key takeaway is pragmatic: LLMs can model heterogeneous schemas well, but they bring new risks around invalid rows, numerical fidelity, and privacy leakage. - [`06_llm_tabular_great`](06_llm_tabular_great.ipynb) — This notebook implements GReaT (Generate Realistic Tabular Data) using the actual be-great library to generate synthetic financial tabular data with LLMs. Uses etf_tabular_data, etfs, from_dir data. ### 5.8 Applying the Fidelity-Utility-Privacy Framework This is the chapter's methodological center of gravity. It argues that validation matters more than architecture and organizes evaluation around fidelity, utility, and privacy, with TSTR as the main task-based benchmark and leakage or DP checks as privacy controls. This section matters because it turns synthetic data from a modeling curiosity into something readers can govern and assess in a research workflow. - [`07_dp_gan`](07_dp_gan.ipynb) — This notebook demonstrates Differential Privacy (DP) for training generative models using Opacus, PyTorch's official DP library. We train a GAN with formal privacy guarantees that limit information leakage about individual records. ## Running the Notebooks ```bash # From the repository root uv run python 05_synthetic_data/<notebook>.py # Test mode (reduced data via Papermill) uv run pytest tests/test_chapter_notebooks.py -v -k "05_synthetic_data" ``` > Runtime callouts (cold-start, no cached checkpoint): > - `01_timegan`: ~12 min, GPU recommended > - `02_tailgan_tail_risk`: ~6 min cold / <30 s with cached checkpoint, GPU recommended > - `04_gtgan_irregular`: ~6 min, GPU recommended > - `05_diffusion_ts`: ~5 min, GPU recommended > - `06_llm_tabular_great`: ~13 min, GPU required (distilgpt2 fine-tune) > - `07_dp_gan`: ~8 min, GPU recommended > `03_sigcwgan_signatures` requires the `py312` docker profile (signatory/esig are x86-only): > ```bash > docker compose --profile py312 run --rm py312 python 05_synthetic_data/03_sigcwgan_signatures.py > ``` ## References - **Martín Abadi et al.** (2016). [Deep Learning with Differential Privacy](https://doi.org/10.1145/2976749.2978318). - **David H. Bailey and Marcos Lopez de Prado** (2014). [The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality](https://doi.org/10.2139/ssrn.2460551). - **David H. Bailey et al.** (2015). [The Probability of Backtest Overfitting](https://doi.org/10.2139/ssrn.2326253). - **Tim Bollerslev** (1986). [Generalized autoregressive conditional heteroskedasticity](https://doi.org/10.1016/0304-4076(86)90063-1). *Journal of Econometrics*. - **Vadim Borisov et al.** (2023). [Language Models are Realistic Tabular Data Generators](https://doi.org/10.48550/arXiv.2210.06280). - **Adil Rengim Cetingoz and Charles-Albert Lehalle** (2025). [Synthetic Data for Portfolios: A Throw of the Dice Will Never Abolish Chance](https://arxiv.org/abs/2501.03993v5). - **R. Cont** (2001). [Empirical properties of asset returns: stylized facts and statistical issues](https://doi.org/10.1080/713665670). *Quantitative Finance*. - **Rama Cont et al.** (2025). [Tail-GAN: Learning to Simulate Tail Risk Scenarios](https://doi.org/10.48550/arXiv.2203.01664). - **Prafulla Dhariwal and Alex Nichol** (2021). [Diffusion Models Beat GANs on Image Synthesis](https://arxiv.org/abs/2105.05233v4). - **Florian Eckerli and Joerg Osterrieder** (2021). [Generative Adversarial Networks in finance: an overview](https://doi.org/10.48550/arXiv.2106.06364). - **Xuefeng Gao et al.** (2025). [Factor-Based Conditional Diffusion Model for Portfolio Optimization](https://doi.org/10.48550/arXiv.2509.22088). - **Jonathan Ho and Tim Salimans** (2022). [Classifier-Free Diffusion Guidance](https://doi.org/10.48550/arXiv.2207.12598). - **Jonathan Ho et al.** (2020). [Denoising Diffusion Probabilistic Models](http://arxiv.org/abs/2006.11239). *arXiv:2006.11239 [cs, stat]*. - **Jinsung Jeon et al.** (2022). [GT-GAN: General Purpose Time Series Synthesis with Generative Adversarial Networks](https://proceedings.neurips.cc/paper_files/paper/2022/hash/f03ce573aa8bce26f77b76f1cb9ee979-Abstract-Conference.html). *Advances in Neural Information Processing Systems*. - **Sohyeon Kwon and Yongjae Lee** (2024). [Can GANs Learn the Stylized Facts of Financial Time Series?](https://doi.org/10.48550/arXiv.2410.09850). - **Hao Ni et al.** (2020). [Conditional Sig-Wasserstein GANs for Time Series Generation](http://arxiv.org/abs/2006.05421). *arXiv:2006.05421 [cs, stat]*. - **Dimitris N. Politis and Joseph P. Romano** (1994). [The Stationary Bootstrap](https://doi.org/10.2307/2290993). *Journal of the American Statistical Association*. - **Marcos López de Prado** (2022). [Type I and Type II Errors of the Sharpe Ratio under Multiple Testing](https://doi.org/10.3905/jpm.2022.1.403). *The Journal of Portfolio Management*. - **Tomonori Takahashi and Takayuki Mizuno** (2024). [Generation of synthetic financial time series by diffusion models](https://doi.org/10.48550/arXiv.2410.18897). - **Jinsung Yoon et al.** (2019). [Time-series Generative Adversarial Networks](http://papers.nips.cc/paper/8789-time-series-generative-adversarial-networks.pdf). *Curran Associates, Inc.*. - **Xinyu Yuan and Yan Qiao** (2024). [Diffusion-TS: Interpretable Diffusion for General Time Series Generation](https://arxiv.org/abs/2403.01742v3).
Reproduit dans son intégralité avec attribution, conformément à la licence de la source. Licence: MIT
Ce résumé a été rédigé par l’agent de recherche de Stratmill à partir de la source originale ; il n’en est pas une copie.