Surrogate and Correlated Time Series for Strategy Testing
Summary
The document surveys ways to create synthetic market series for testing trading strategies, motivated by daily bond-price and credit default swap index data. One proposed technique comes from surrogate-data methods: transform a series into the frequency domain, randomize its phases, and apply the inverse transform. The aim is to preserve some basic statistical features while changing aspects of the original data. A separate suggestion is to generate correlated series using correlations estimated from market observations.
The responses point to examples and references, including an FFT-based price-generation example, and report that one example showed statistical similarity between original and generated series. These are suggestions rather than a comparison of methods or a formal validation. A central caveat is that surrogate methods may not preserve serial dependence, which can matter for a strategy. Matching observed correlations alone also does not ensure realistic dynamics or dependence in the tails. The document offers no specific diagnostics for deciding whether a synthetic dataset is suitable for a given test.
Key ideas
- Surrogate methods can create altered series that retain selected statistical properties of observed data.
- One suggested approach randomizes Fourier-domain phases before transforming the series back to the time domain.
- Correlated synthetic series can be generated using correlations estimated from market data.
- Preserved basic statistics do not guarantee that important serial dependence is retained.
- Synthetic data should be judged against the features that matter to the strategy being tested.
Tags
Full text
# Literature on generating synthetic time series for testing # Literature on generating synthetic time series for testing I have some market data (daily time series) for bond prices and CDS indices and I would like to generate synthetic versions of these which are statistically "similar" for testing trading strategies. Is there literature on this subject? ## Answer by Doodles (score 7) https://quant.stackexchange.com/a/3039 Yes, there is in fact a whole literature on this subject coming from the field of non-linear dynamics-- it is known as the method of surrogates. The idea is essentially to come up with a "scrambled" version of your original data set that preserves many of the basic statistical properties, though perhaps not the serial dependence structure which might be important for your purposes. I think the best way to do this is to apply a Fourier transform to your data to re-express it in the frequency domain, and then randomly shuffle the data in that form, then convert back by applying the inverse transformation. This Matlab code from the FileExchange does exactly that: http://www.mathworks.com/matlabcentral/fileexchange/32621-phase-randomization/content/phaseran.m ## Answer by Lliane (score 4) https://quant.stackexchange.com/a/1841 The easiest answer which comes to my mind is generating correlated time series. There is a good description of the process in Part 3 of this document : http://www.columbia.edu/~mh2078/MCS04/MCS_framework_FEegs.pdf If you base your correlation input on the correlation observed in the market data you should obtain "statistically" similar time series. ## Answer by babelproofreader (score 1) https://quant.stackexchange.com/a/3047 In response to this question, I posted an answer which in turn links to a post on my blog which uses the Fast Fourier Transform to create synthetic prices, in a similar vein to the code linked to in Doodles' answer above. The example screen shots in my blog post show terminal output which indicate high statistical similarity between the original time series and the synthetic series generated from it.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.