Simulating Synthetic Share Prices with GBM, Jump Models, and Random Walks
Summary
The document considers ways to create realistic-looking share price histories for data visualization without distributing a real company’s data. It identifies geometric Brownian motion (GBM) as a common starting model: prices are log-normal, with returns driven by a continuous random process. It also notes that jump models, described as Lévy models, can represent price movements more realistically, with GBM as a special case. A random walk is offered as another simple way to generate an unpredictable path.
The responses do not compare these approaches or provide parameter choices, code, or evidence that any one produces representative market data. The document also suggests using existing datasets from the UCI Machine Learning Repository, subject to its citation policy and any dataset-specific requests. That suggestion concerns obtaining reusable real data rather than generating synthetic prices. The discussion is brief, so it does not address calibration, realistic return distributions, volatility clustering, or the limitations of using simulated series for research.
Key ideas
- Geometric Brownian motion is a standard starting point for simulating share prices.
- Jump models can represent price paths with discontinuous moves, and GBM is a special case of the broader Lévy model family.
- A random walk offers a simple way to create an unpredictable path.
- Reusable real datasets may be an alternative, subject to their citation and usage terms.
- The document gives no calibration guidance or evidence comparing the suggested approaches.
Tags
Full text
# Looking for an algorithm to generate "dummy" share price data # Looking for an algorithm to generate "dummy" share price data Is there an easy-ish way I can generate "dummy" share price data for the purposes of data visualisation techniques etc.? Essentially I want to have something like the "Adventure Works" of price data. I don't want to use actual price data of a real company due to issues over ownership/distribution of the data. But something that looks similar to a "typical" history of share prices (over 1 year, say) in terms of its data points is what I'm after. I considered random deviations from a "start" price, is there a better way? ## Answer by Richi Wa (score 2, accepted) https://quant.stackexchange.com/a/23060 So you want to simulate a price path, right? Googeling this you will find code in the programming language of your choice. The usual starting point for share price simulation is the Geometric Brownian motion (GBM) model. It assumes log-normal prices. It turns out that models with jumps often relfect reality closer. Such models are called Lévy models and the GBM is a special case. That's where you should start. Looking for similar questions here, you find a lot too. E.g. This one. ## Answer by Ogaday (score 2) https://quant.stackexchange.com/a/23059 I don't know enough about share prices to answer your question, but the UCI Machine Learning Repository is an good resource for real data that you'll likely be able to use for any purpose, as long as you cite them. From their website: > The UCI Machine Learning Repository is a collection of databases, domain theories, and data generators that are used by the machine learning community for the empirical analysis of machine learning algorithms. The archive was created as an ftp archive in 1987 by David Aha and fellow graduate students at UC Irvine. Since that time, it has been widely used by students, educators, and researchers all over the world as a primary source of machine learning data sets. As an indication of the impact of the archive, it has been cited over 1000 times, making it one of the top 100 most cited "papers" in all of computer science. The citation policy: > If you publish material based on databases obtained from this repository, then, in your acknowledgements, please note the assistance you received by using this repository. This will help others to obtain the same data sets and replicate your experiments. and > A few data sets have additional citation requests. These requests can be found on the bottom of each data set's web page. If you filter by "business" there are few datasets which look like they might contain the right type of information. It might be worth the cost of citing them in order to use them (only necessary if making your public, I believe) because they are free to use for any purpose otherwise, as far as I can see. ## Answer by Ignorant (score 1) https://quant.stackexchange.com/a/23087 Another fascinating option: the so-called random walk which is, in layman's terms the simulation of a random and unpredictable path. More information on Wikipedia and an implementation experiment in Python.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.