Time-Series Validation for Logistic Return Prediction
Summary
The document addresses whether logistic regression can be used to predict the direction of tomorrow’s return and how to validate such a model. It explains that regression is commonly used for out-of-sample generalization, but randomly assigning observations to training and test sets is unsuitable for time-series data: a model can train on later observations while being evaluated on earlier ones, creating future-information leakage and potentially overstating performance.
It recommends chronological walk-forward validation, with time-series splits or purged folds and an embargo to reduce leakage between folds. After choosing a sound validation design, the challenge is to develop features that carry information about future returns and remain useful across changing market regimes, or to update the model regularly. The document offers methodological guidance rather than empirical results; it does not establish that any particular feature set or logistic model will predict returns successfully.
Key ideas
- Logistic regression can be used to make out-of-sample predictions.
- Randomly splitting time-series observations can leak future information into training.
- Walk-forward validation preserves the time order of observations.
- Purged folds with an embargo can help limit leakage between validation folds.
- Predictive features must be robust to market-regime changes or the model may need regular updates.
Tags
Full text
# Using regression for binomial prediction of tomorrow's return # Using regression for binomial prediction of tomorrow's return I completed a challenge which asks the user to predict tomorrow's market return. The data available is prices data and the model must be logistic regression. They call it "machine learning" but it's pure regression. I derived a number of features from the data (e.g. momentum) etc. The challenge implies that I need to train a model to predict returns out-of-sample. This is "extrapolation" which is not what regression is usually used for. Is this a valid use of regression? I split my train/test data into 80%/20%. The split is performed by random sampling. What I mean is that for example, if I have data for 2010-2020, I don't use 2010-2018 for training and 2018-2020 for test but rather I randomly picked elements from the data such that 80% is train. This means that 2010-01-05 might belong to the train dataset but 2010-01-06 (next day) might belong to the test dataset. My questions are: - Are my points about the use of regression here valid? - Does my test/train split mitigate the extrapolation problem with regression? - Does the random sampling of train/test cause problems due to the timeseries nature of the data? ## Answer by autoencoder (score 2) https://quant.stackexchange.com/a/71673 (1) Regression is classic machine learning, and the general goal of ML is to train a model that can generalize to unseen (out of sample) data. Also, in the context of your challenge and in practice, people indeed care more about predicting the future (extrapolation) than fitting perfectly what has happened in the past, so using regression for this is valid. (2, 3) Random split will not solve your problem, it will directly leak future information to your model and thus increase overfitting (returns can have high autocorrelation), due to the timeseries nature of the data. A better cross validation setup is to use walk-forward data split. You could, for example, use TimeSeriesSplit or Purged K-FOLD cross validation with embargo to avoid leak between folds. After you have the right validation setup, your main problem is to design features that can really explain future returns and are robust enough to go through market regime changes, or update your model regularly.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.