Skip to content
All library documents

Modeling Stochastic Data with Time-Series Methods Before SDEs

Article Quant Q&A · Author: Kilkik

Summary

The discussion asks how to infer a stochastic process from a long observed series and questions whether a histogram can identify its dynamics. The response recommends beginning with discrete-time time-series analysis rather than trying to recover an underlying stochastic differential equation. It suggests checking whether the series is bounded, considering logarithms for one-sided bounded data, and using an automated ARIMA selection process to examine integration and stationarity under a homoscedasticity assumption.

If trends or seasonality drive nonstationarity, those features should be addressed before fitting the time-series model. If constant variance is not plausible, the response suggests an ARMA–GARCH approach; residuals may then be examined or given a parametric distribution if needed. Another answer proposes testing the log series or level series for normality, or considering a mixture model estimated by likelihood. These are practical starting points, not a universal identification procedure: model choice depends on the goal, and the discussion gives no data, diagnostics, or fitted results to establish which model applies.

Key ideas

  • A stochastic process may be modeled adequately in discrete time without identifying its SDE.
  • Check boundedness and consider transforming one-sided bounded data with logarithms.
  • ARIMA methods can help assess integration and stationarity when variance is approximately constant.
  • Address trends or seasonality before fitting stationary dynamics, and consider ARMA–GARCH when variance changes.
  • Distribution tests and mixture models are possibilities, but the modeling choice depends on the task.

Tags

Full text
# Identifying stochastic process from data


# Identifying stochastic process from data












Suppose I am given the values of a stochastic process $S_t$ satisfying some unknown SDE from say 2000 to 2024 so I have a lot a data. How do I identify, model this stochastic process ?

First I thought of making a histogram of its values and try to guess what distribution it comes from but histograms are made to guess the distribution of i.i.d. samples meanwhile here each $S_t$ follows is probably $f(W_t$) where $W_t$ is a brownian motion so a histogram isn't useful.

Is there any way do it because I never heard of such type of modeling. Also, I intuitively don't think there are ways since two geometric brownian motions with same parameters can be so different for example, so how to make a good guess ?

## Answer by achirikhin (score 1)

https://quant.stackexchange.com/a/79520

A model that cannot be rejected for the given observations does not need to be described by an SDE. It may suffice to stay within the "discrete time" time series analysis.

- See if your time series is bounded. If it is one-side bounded, it may makes sense to model the logarithm.

- Assuming homoschedasticity (no "local vol" in the SDE terms), try using some Auto-ARIMA tool to estimate such ARIMA. This will firstly pick up integration, e.g. stock value vs its return (if you have not done this yet by switching to logs) and then show if the process is at least stationary. If such ARIMA is estimated, then you are done. All you need is to extract the residuals and, if really necessary, estimate a parametric model for them, which may turn out to be Gaussian indeed.

- If 2 does not work, then you will first have to identify the source of non-stationarity beyond integration (trends, seasonality), correct for those and goto 2.

- In rare occasions when homoscedasticity is rejected, you may have to resort to ARMA-GARCH

This stuff is more practical and far less restrictive than trying to identify an SDE.

Good luck

## Answer by Arshdeep (score 0)

https://quant.stackexchange.com/a/79155

To see if it's lognormal, check if the log follows a normal distribution (statistical test, there are many). To see if it is normal, check if itself follows a normal distribution. Maybe you want a mixture and calibrate the parameters using max likelihood. It ALL depends on your end goal - what do you want to do?

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.