Using Time-Series GANs to Create Synthetic Data for Strategy Backtests
Summary
The article explains how a time-series generative adversarial network can produce synthetic financial observations when historical data is limited. It describes the generator and discriminator conceptually, then focuses on the conditional probabilistic autoregressive synthesizer in the Synthetic Data Vault. This approach is designed for multiple asset sequences with stable context fields; it cannot be applied to a single asset series on its own.
The example uses Apple and Microsoft data to generate synthetic volume and returns, reconstruct prices, and augment inputs for a random forest strategy. It describes a walk-forward process that repeatedly fits models and selects among them using a Sharpe ratio on test data. The article presents implementation choices and caveats rather than evidence that synthetic data improves live performance. Synthetic samples can reflect limitations or artifacts in the source data, and the described price reconstruction omits high and low inputs to avoid inconsistent OHLC values. The author also recommends multiple generated paths and transaction costs, and notes the compute tradeoff from training epochs.
Key ideas
- GANs train a generator and discriminator in competition to create data resembling a training set.
- The described synthesizer requires multiple time series and stable context information for each sequence.
- The example augments real stock history with synthetic volume and return data for a random forest backtest.
- The walk-forward process compares models using test-period Sharpe ratios, but does not establish live profitability.
- More training epochs may improve generated samples while increasing fitting time, and transaction costs should be included.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.