Pular para o conteúdo
Todos os documentos da biblioteca

Um universo temático de ETF para pesquisa de momentum

Artigo Machine Learning for Trading

Resumo

Esta nota de dados descreve uma coleção diversificada de fundos negociados em bolsa, usada em uma estratégia de momentum e em uma série mais ampla de exemplos de pesquisa financeira. Ela fornece observações diárias de abertura, máxima, mínima, fechamento e volume a partir de 2006, agrupadas entre estilos e setores de ações dos US, mercados internacionais e emergentes, renda fixa, commodities, temas especializados e moedas. O documento também explica a origem, a organização e os metadados do conjunto de dados, além das opções de carregamento para selecionar símbolos e intervalos de datas.

A diversidade dos grupos de ativos torna a coleção adequada para pesquisa transversal e multiclasse, e seus usos posteriores declarados incluem desenvolvimento de estratégias, análise de atributos, estudos de robustez, execução e relatórios de desempenho. A nota não relata retornos de uma regra de negociação nem comprova que os dados sustentem uma estratégia lucrativa. Os dados vêm do Yahoo Finance sob os termos pessoais e educacionais declarados; a redistribuição dos dados brutos é restrita, e a cobertura do provedor, os ajustes e a qualidade dos dados devem ser considerados em conclusões empíricas.

Ideias principais

  • A coleção contém histórico diário de OHLCV para 100 ETFs em nove grupos temáticos.
  • A cobertura abrange exposições a ações, renda fixa, commodities, moedas e mercados internacionais.
  • Um esquema comum e os metadados de símbolos permitem carregar subconjuntos por ticker e intervalo de datas.
  • O conjunto de dados é usado em pesquisas de momentum e em vários exemplos de análise e execução.
  • Os termos declarados pela fonte restringem a redistribuição das observações brutas.

Tags

Texto completo
# Chapter 5: Synthetic Financial Data


# Chapter 5: Synthetic Financial Data

Financial research has one realized price path and a research process that searches it repeatedly, which is enough on its own to make a backtest fragile. This chapter asks what it would take to generate more paths worth trusting: what a generator has to preserve, how to tell whether it did, and what the answer costs in privacy and in effort. It works through classical simulation baselines, GANs, diffusion models and LLM-based tabular generation, and it holds each of them to the same evaluation, because the validation protocol matters more here than the choice of architecture.

## Learning Objectives

* Explain why trading research is path-limited and how adaptive search and multiple testing can inflate apparent backtest performance.
* Use classical simulation baselines, including bootstrap and stochastic volatility models, as interpretable benchmarks for synthetic data generation.
* Select a synthetic-data approach that matches the data structure and downstream objective, including learned generators for time series and tabular financial data.
* Diagnose generated data using stylized-fact, dependence, and task-based evaluation methods, including Train-Synthetic-Test-Real comparisons.
* Assess privacy and generator-specific risks, including leakage, bias amplification, overfitting to the generator, and limited scenario novelty.

## Sections

### 5.1 The Quant's Dilemma

This section explains why backtests are fragile even before any generative model enters the picture. Because strategy research is adaptive and path-limited, strong in-sample results may simply reflect favorable history and repeated search. The section then motivates synthetic data as a way to expand robustness analysis beyond one realized market path, while anchoring the discussion in the stylized facts that any useful generator must preserve.

### 5.2 Evaluating Synthetic Financial Data

This section sets the standard the rest of the chapter is measured against, which is why it comes before any generator. Synthetic data can look plausible and still be unusable, because the structure that matters in finance sits in rare events, in shifts in dependence, and in conditional dynamics rather than at the centre of the distribution. It introduces the three criteria the chapter uses throughout - fidelity, utility and privacy - and the fact that they trade off against one another rather than improving together.

### 5.3 Classical Simulation Baselines

This section gives the chapter intellectual discipline. Rather than jumping straight to deep generators, it shows that bootstrap methods, GBM, jump-diffusion, mean reversion, Heston, and GARCH remain important baselines because they are interpretable, sample-efficient, and easier to validate. Readers should care because these methods define the benchmark a learned generator ought to beat on the diagnostics that actually matter.

- [`00_classical_simulation`](00_classical_simulation.ipynb) — Simulates return paths from GBM, jump diffusion, mean reversion, Heston and GARCH, and resamples real returns with the IID, block and stationary bootstraps, then measures which of them reproduce fat tails and volatility clustering. These are the baselines a learned generator has to beat. Uses ETF data.

### 5.4 Generative Model Taxonomy

This section provides the conceptual map for the rest of the chapter. It distinguishes discriminative from generative modeling and positions VAEs, GANs, diffusion models, and LLM-based tabular generators as alternative ways to learn joint distributions rather than hand-specify them. Its value is orientation: readers can see early that architecture choice is really about preserving the structure their downstream use case needs.

### 5.5 GANs for Financial Time Series

This is the chapter's most differentiated model survey. It moves from the basic adversarial setup to finance-specific variants such as TimeGAN, Tail-GAN, Sig-CWGAN, and GT-GAN, showing how each responds to a concrete weakness of vanilla GANs: temporal structure, tail risk, path fidelity, or irregular timestamps. The section matters because it teaches readers not to ask which GAN is best in general, but which inductive bias matches the task and its failure modes.

- [`01_timegan`](01_timegan.ipynb) — This notebook implements TimeGAN (Yoon, Jarrett & van der Schaar, NeurIPS 2019), the foundational architecture for synthetic financial time series generation. Uses multi_stock_data, state_dict, us_equities data.
- [`02_tailgan_tail_risk`](02_tailgan_tail_risk.ipynb) — This notebook implements Tail-GAN (Cont, Xu, and Zhang 2022), a GAN architecture that uses differentiable sorting to preserve tail risk characteristics (VaR, ES) in synthetic financial scenarios. Uses etf_returns, etfs, state_dict data.
- [`03_sigcwgan_signatures`](03_sigcwgan_signatures.ipynb) — Builds the unconditional Sig-Wasserstein GAN, which replaces the learned discriminator with an analytic distance between expected path signatures, and explains how the conditional Sig-CWGAN extends it. Uses S&P 500 log returns. Needs the `py312` profile; see the callout at the end of this file.
- [`04_gtgan_irregular`](04_gtgan_irregular.ipynb) — This notebook implements a GT-GAN-inspired model (based on Jeon et al., NeurIPS 2022) using Neural ODEs to handle time series with naturally irregular timestamps. Uses Chapter 3 NVDA dollar bars (Databento bar sampling).

### 5.6 Diffusion Models for Financial Time Series

This section presents diffusion models as a strong, often more stable alternative to adversarial training. It explains the denoising framework, why diffusion can fit financial return structure, how conditional guidance supports regime-aware stress testing, and why Diffusion-TS is a useful reference design for sequential financial data. Readers should care because this is the chapter's clearest candidate for a general-purpose learned generator, but one that still demands careful validation and use-case alignment.

- [`05_diffusion_ts`](05_diffusion_ts.ipynb) — This notebook implements Diffusion-TS (Yuan & Qiao, ICLR 2024), a diffusion model that decomposes the denoising prediction into trend (polynomial regression) and seasonal (Fourier basis) components. This interpretable structure encourages the model to separate slow drift from periodic patterns, analogous to classical STL decomposition but learned end-to-end within the diffusion framework.

### 5.7 LLMs for Structured Financial Data

This section extends the synthetic-data discussion beyond return series to mixed-type financial tables. By introducing serialization, the GReaT workflow, and constraint-based postprocessing, it shows where LLMs can be practical for credit, customer, and fundamental data. The key takeaway is pragmatic: LLMs can model heterogeneous schemas well, but they bring new risks around invalid rows, numerical fidelity, and privacy leakage.

- [`06_llm_tabular_great`](06_llm_tabular_great.ipynb) — This notebook implements GReaT (Generate Realistic Tabular Data) using the actual be-great library to generate synthetic financial tabular data with LLMs. Uses etf_tabular_data, etfs, from_dir data.

### 5.8 Applying the Fidelity-Utility-Privacy Framework

This is the chapter's methodological center of gravity. It argues that validation matters more than architecture and organizes evaluation around fidelity, utility, and privacy, with TSTR as the main task-based benchmark and leakage or DP checks as privacy controls. This section matters because it turns synthetic data from a modeling curiosity into something readers can govern and assess in a research workflow.

- [`07_dp_gan`](07_dp_gan.ipynb) — This notebook demonstrates Differential Privacy (DP) for training generative models using Opacus, PyTorch's official DP library. We train a GAN with formal privacy guarantees that limit information leakage about individual records.

## Running the Notebooks

```bash
# From the repository root
uv run python 05_synthetic_data/<notebook>.py

# Test mode (reduced data via Papermill)
uv run pytest tests/test_chapter_notebooks.py -v -k "05_synthetic_data"
```

> Runtime callouts (cold-start, no cached checkpoint):
> - `01_timegan`: ~12 min, GPU recommended
> - `02_tailgan_tail_risk`: ~6 min cold / <30 s with cached checkpoint, GPU recommended
> - `04_gtgan_irregular`: ~6 min, GPU recommended
> - `05_diffusion_ts`: ~5 min, GPU recommended
> - `06_llm_tabular_great`: ~13 min, GPU required (distilgpt2 fine-tune)
> - `07_dp_gan`: ~8 min, GPU recommended

> `03_sigcwgan_signatures` requires the `py312` docker profile (signatory/esig are x86-only):
> ```bash
> docker compose --profile py312 run --rm py312 python 05_synthetic_data/03_sigcwgan_signatures.py
> ```

## References

- **Martín Abadi et al.** (2016). [Deep Learning with Differential Privacy](https://doi.org/10.1145/2976749.2978318).
- **David H. Bailey and Marcos Lopez de Prado** (2014). [The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality](https://doi.org/10.2139/ssrn.2460551).
- **David H. Bailey et al.** (2015). [The Probability of Backtest Overfitting](https://doi.org/10.2139/ssrn.2326253).
- **Tim Bollerslev** (1986). [Generalized autoregressive conditional heteroskedasticity](https://doi.org/10.1016/0304-4076(86)90063-1). *Journal of Econometrics*.
- **Vadim Borisov et al.** (2023). [Language Models are Realistic Tabular Data Generators](https://doi.org/10.48550/arXiv.2210.06280).
- **Adil Rengim Cetingoz and Charles-Albert Lehalle** (2025). [Synthetic Data for Portfolios: A Throw of the Dice Will Never Abolish Chance](https://arxiv.org/abs/2501.03993v5).
- **R. Cont** (2001). [Empirical properties of asset returns: stylized facts and statistical issues](https://doi.org/10.1080/713665670). *Quantitative Finance*.
- **Rama Cont et al.** (2025). [Tail-GAN: Learning to Simulate Tail Risk Scenarios](https://doi.org/10.48550/arXiv.2203.01664).
- **Prafulla Dhariwal and Alex Nichol** (2021). [Diffusion Models Beat GANs on Image Synthesis](https://arxiv.org/abs/2105.05233v4).
- **Florian Eckerli and Joerg Osterrieder** (2021). [Generative Adversarial Networks in finance: an overview](https://doi.org/10.48550/arXiv.2106.06364).
- **Xuefeng Gao et al.** (2025). [Factor-Based Conditional Diffusion Model for Portfolio Optimization](https://doi.org/10.48550/arXiv.2509.22088).
- **Jonathan Ho and Tim Salimans** (2022). [Classifier-Free Diffusion Guidance](https://doi.org/10.48550/arXiv.2207.12598).
- **Jonathan Ho et al.** (2020). [Denoising Diffusion Probabilistic Models](http://arxiv.org/abs/2006.11239). *arXiv:2006.11239 [cs, stat]*.
- **Jinsung Jeon et al.** (2022). [GT-GAN: General Purpose Time Series Synthesis with Generative Adversarial Networks](https://proceedings.neurips.cc/paper_files/paper/2022/hash/f03ce573aa8bce26f77b76f1cb9ee979-Abstract-Conference.html). *Advances in Neural Information Processing Systems*.
- **Sohyeon Kwon and Yongjae Lee** (2024). [Can GANs Learn the Stylized Facts of Financial Time Series?](https://doi.org/10.48550/arXiv.2410.09850).
- **Hao Ni et al.** (2020). [Conditional Sig-Wasserstein GANs for Time Series Generation](http://arxiv.org/abs/2006.05421). *arXiv:2006.05421 [cs, stat]*.
- **Dimitris N. Politis and Joseph P. Romano** (1994). [The Stationary Bootstrap](https://doi.org/10.2307/2290993). *Journal of the American Statistical Association*.
- **Marcos López de Prado** (2022). [Type I and Type II Errors of the Sharpe Ratio under Multiple Testing](https://doi.org/10.3905/jpm.2022.1.403). *The Journal of Portfolio Management*.
- **Tomonori Takahashi and Takayuki Mizuno** (2024). [Generation of synthetic financial time series by diffusion models](https://doi.org/10.48550/arXiv.2410.18897).
- **Jinsung Yoon et al.** (2019). [Time-series Generative Adversarial Networks](http://papers.nips.cc/paper/8789-time-series-generative-adversarial-networks.pdf). *Curran Associates, Inc.*.
- **Xinyu Yuan and Yan Qiao** (2024). [Diffusion-TS: Interpretable Diffusion for General Time Series Generation](https://arxiv.org/abs/2403.01742v3).

Exibido na íntegra, com atribuição conforme a licença da fonte. Licença: MIT

Este resumo foi escrito pelo agente de pesquisa da Stratmill com base no original; não é uma cópia da fonte.