Skip to content
All library documents

Avoiding Leakage and Regime Assumptions in ML Time-Series Forecasting

Article Quant Q&A · Author: Pur

Summary

The document describes a workflow for forecasting financial time series with technical indicators and machine learning: calculate indicators, remove highly correlated features, lag inputs, train with a rolling or expanding window, and evaluate on held-out data. The questioner reports overly optimistic results and asks whether indicators should be calculated separately for training and test periods. The replies caution that a model may rely on relationships that do not persist, since financial patterns and relevant historical regimes can change over time.

The responses suggest considering classification of extreme returns, a VARMA model, or an echo state network, and emphasize recent data and prediction uncertainty. They do not directly resolve the indicator-calculation question or diagnose the reported optimism. In practice, indicator calculations should use only information available at each forecast date, and feature selection and other fitted preprocessing should be confined to the training data within each time-series validation split. The document supplies no empirical comparison or validated forecasting results.

Key ideas

  • The proposed workflow uses lagged technical indicators, correlation-based feature filtering, rolling training, and held-out evaluation.
  • Financial relationships may shift over time, making stable-parameter assumptions questionable.
  • The replies mention extreme-return classification, VARMA, and echo state networks as possible approaches.
  • Confidence intervals or predictive distributions can help assess uncertainty in forecasts.
  • The document does not diagnose the optimistic results or fully answer the indicator timing question.

Tags

Full text
# Using Technical Indicators for forecasting Financial time series using Machine learning models


# Using Technical Indicators for forecasting Financial time series using Machine learning models












Hi I am trying to use financial technical Indicators for forecasting, using machine learning models. The usual approach in time series cross validation is to use a moving window or growing window. The methodology I am using is described in the following steps

- Calculate technical indicators TA1, TA2, ....TAN for the whole historical data set, using lag 1

- Then use simple feature selection method like finding out the cross correlation between the independent variables, and remove variables with cross correlation above a certain threshold

- Lag the input variables by one

- Then train the the Machine learning model using a moving window, then reports its performance on the train set

- Test its performance on a test set which was not used in the training process

The issue I am facing is that the results are highly optimistic. My question is should I calculate the technical indicators separately for the train and test and then use them, or should I calculate them for the whole data set in the beginning and then divide them into train and test set.

## Answer by Kyle Balkissoon (score 1)

https://quant.stackexchange.com/a/16000

In terms of forecasting, it is VERY difficult to forecast financial time series especially using ML models. One of the "successful" papers that I have seen use a classifier approach (e.g. forecasting extreme returns).

See: http://algorithmicfinance.org/2-1/pp45-58/

The above being said, your model structure would assume that the parameters are stable across time. Why not use a VARMA approach using the TA indicators see here: http://users.monash.edu.au/~gathana/slides/isf07.pdf

## Answer by user3001408 (score 1)

https://quant.stackexchange.com/a/16004

Careful when you use machine learning techniques in financial time series. You are implicitly assuming that the trend that you spot on the data you train your model on, the same trend will be there in future time series.

In addition to what KKB suggested, there is another model called "Echo State Network" (a family of RNNs) you should look at. But then again you are advised to train on brief data. Reason being - what happened 2 years back may not be relevant now, but what's happening in past few months are relevant as they reflect recent events. It's a maze really. You have to use your intuition.

Also use the confidence interval to judge how meaningful your estimated value is. A probability distribution is quite helpful in gauging the effectiveness of the prediction.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.