Skip to content
All library documents

Synthetic FX Data: Model Conditional Volatility and Tick Activity

Article Quant Q&A · Author: JeremyKun

Summary

The document discusses generating synthetic foreign-exchange data for algorithm testing, including return paths, OHLC bars, tick timing, and bid-ask spreads. Its main caution is that Gaussian noise around a trend is a poor representation of high-frequency FX returns, which can have fat tails, changing volatility, and time-of-day or weekly patterns.

For simulations intended to study volatility behavior or system robustness under changed data-generating conditions, it recommends modeling conditional heteroskedasticity and tick frequency, then constructing bars from simulated ticks. It also mentions bootstrapping representative historical observations and data augmentation as possible approaches. Synthetic results should not be treated as reliable evidence for accepting or rejecting strategies without a realistic model: the suitable method depends on the properties the simulation is meant to preserve, and the discussion offers no single validated recipe for spreads or tick intervals.

Key ideas

  • Gaussian perturbations are a weak model for high-frequency FX returns because return distributions can be fat-tailed and volatility can vary over time.
  • Conditional heteroskedastic models can represent changing volatility better than one unconditional distribution.
  • OHLC bars can be built from simulated tick paths using a separate tick-arrival process.
  • Historical bootstrapping and data augmentation are possible alternatives, but depend on which data features must be preserved.
  • Synthetic backtest results may mislead when the simulated process does not reflect the intended market behavior.

Tags

Full text
# How to generate synthetic FX data for backtesting?


# How to generate synthetic FX data for backtesting?












I want to generate synthetic forex data for the purpose of backtesting my trading algorithms. I have some rough ideas in mind on how to do this:

> Start with a curve representing a trend, then randomly generate points around the curve according to a Gaussian or some other distribution. Then take the generated points and somehow generate the bar data (open, high, low, close) around those points; alternatively, add a time factor, and then randomly determine when a tick occurs and collect the data into bars afterward.

My question is: is this far off from the established methods for synthetic data generation? I suppose that raises the more basic question: are there any established methods for synthetic data generation? I can't seem to find any writing on this subject, be it a blog post or a research paper.

So in addition to a request for external resources on synthetic data generation, I'd like to know what sorts of distributions best model the relationships between open, high, low, close, or how to generate the appropriate intervals between ticks, the spread between ask and bid prices, etc.

## Answer by Tal Fishman (score 7)

https://quant.stackexchange.com/a/1689

For starters, I am not even sure why you need to ask this question. There is literally years of free tick data available for FX, just check out quant.SE's data wiki.

Having said that, a Gaussian is a very poor fit to high-frequency data, particularly FX. Your strategy for simulating data depends on the idea behind the simulation. If you wish to actually estimate any parameters or test methods on this data, and you will be accepting or rejecting ideas based on the results of the simulation, I urge you to stop right there and reconsider.

If, instead, you wish to model the volatility of the process and test the ability of your system to deal with changes to the parameters of the data generating process, you should consult chapter 5 and particularly page 122 of An Introduction to High Frequency Finance. They write:

> The distributions of returns are increasingly fat-tailed as data frequency increases (smaller interval sizes) and are hence distinctly unstable... Scaling laws describe mean absolute returns and mean squared returns as functions of their time intervals... There is evidence of seasonal heteroskedasticity in the form of distinct daily and weekly clusters of volatility... Some papers claim FX returns to be close to Paretian stable ones, for instance (McFarland et al., 1982; Westerfield, 1997); some to Student distributions that are not stable (Rogalski and Vinso, 1978; Boothe and Glassman, 1987); some reject any single distribution... Most researchers now agree that a better description of the data generating process is in the form of a conditional heteroskedastic model rather than being from an unconditional distribution.

Finally, if you wish to construct OHLC data, your best bet is to model the data generating process as above along with a tick frequency process (also discussed in the above reference) and construct OHLC bars from simulated tick data.

## Answer by babelproofreader (score 1)

https://quant.stackexchange.com/a/1692

The method for generating synthetic data described here might be useful to you. Also I believe the meboot R package can be used for synthetic time series generation.

## Answer by G__ (score 0)

https://quant.stackexchange.com/a/1705

You could start with some representative data (i.e. historical data for the period of interest) and then use bootstrapping to estimate the true distribution of that data. From there, you can use that distribution to generate representative data.

- In R

- In Incanter

## Answer by Peter Cotton (score 0)

https://quant.stackexchange.com/a/29668

I've been looking at the same question. I believe the data augmentation literature is relevant. And no doubt some ideas from Good-Turing frequency estimation and descendants can be adapted. Another idea is to transform one exchange rate so that it roughly coincides with another, and use a barrage of off-the-shelf classification algorithms to see if they can determine fake from real data. I think it all depends on which invariants in the data you believe you know well and which residual behavior is common across different time series.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.