Stationarity Considerations for Time-Series Prediction and Clustering
Summary
The document discusses whether time-series features must be weakly stationary for neural networks, other predictive machine learning models, or clustering. It argues that nonstationary predictive inputs can encourage a model to rely mainly on autocorrelation, potentially reducing a one-step forecast to dependence on the latest observation. It also notes that trend and seasonal structure may be modeled explicitly, while detrending or deseasonalizing can sometimes improve forecast accuracy.
The answer distinguishes forecasting from clustering: stationarity is presented as important for predictive time-series modeling, whereas clustering can group observations by feature-space similarity without requiring stationarity. For temporal clustering, stationarity may still help preprocessing or accuracy. The evidence cited is a brief reference to prior research on preprocessing and forecasting errors, but no results or conditions are detailed. The guidance is therefore a general modeling consideration rather than a universal requirement; suitable preprocessing depends on the data, model, and forecasting objective.
Key ideas
- Nonstationary time-series inputs may cause predictive models to lean heavily on autocorrelation.
- Trend and seasonal components can be modeled, but preprocessing them may improve forecasts.
- Clustering does not inherently require stationary data because it groups observations by feature similarity.
- Stationarity can still be useful during temporal clustering preparation.
Tags
Full text
# Weak Stationarity for Neural Network Input?
# Weak Stationarity for Neural Network Input?
I am taking a course that detailed that input data into neural networks should be at least weakly predictive and weakly stationary (stable mean).
Does this principle apply to other ML models like tree-based algorithms or clustering?
Moreover, in general, is stationarity of features always required to construct meaningful time-series analysis?
## Answer by Pleb (score 2, accepted)
https://quant.stackexchange.com/a/79742
Yes, it is essential to ensure that your time-series exhibit atleast some form of stationarity for predictive time-series based ML models.
When estimating a Neural Network (NN) or other machine learning models on non-stationary time-series data, the model may primarily capture the auto-correlation present in the data. Consequently, for a one-step ahead prediction of $X_{t+1}|\mathcal{F}_t$, the model degenerates to using only $X_t$ as the predictor. This is undesirable, as the model fails to utilize the richness of the input data to generate meaningful predictions. A brief LinkedIn Article explains this issue.
It is important to note that you can use trend- and/or seasonal-stationary input-data for your NN models. Depending on the complexity of the model, the NN model will try to capture the trend + seasonality of the time-series and use these as separate predictors. However, research by G.P Zhang et al. (2005) do suggest that de-trending and/or de-seasonalization can significantly reduce forecasting errors for your model, as opposed to using unprocessed raw data.
Regarding clustering, there is no requirement for a time-series to be stationary, as clustering models can group data points based on their similarities or distances in the feature space at a given time $t$. Even with temporal clustering, while stationarity is not a strict requirement, it can aid in data pre-processing and, in some cases, enhance clustering accuracy.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.