Skip to content
All library documents

Regime Shifts as a Cause of Failed Model Generalization

Article Quant Q&A · Author: Vladimir Belik

Summary

The question describes a balanced financial time-series classification task in which several machine-learning algorithms perform well under time-ordered cross-validation but fail on a later test set. The response proposes that the underlying relationship may not come from one stable distribution. Instead, markets may pass through regimes in which relationships, trading rules, or prevailing conditions differ.

Under this explanation, cross-validation can look strong when its training and validation periods share similar regimes, yet offer little evidence about performance in a distinct future regime. The answer gives examples of changing market narratives involving China, risk appetite, and liquidity, but supplies no measured results or method for identifying regime boundaries. Regime change is a plausible diagnosis rather than a demonstrated cause in the questioner’s data. The discussion also does not prescribe a particular validation scheme or modeling remedy, so researchers would need additional analysis to distinguish regime shifts from leakage, tuning bias, or ordinary sampling variation.

Key ideas

  • Strong time-ordered cross-validation results may not predict performance in a future market regime.
  • Financial relationships can change as economic conditions and trading norms shift.
  • A model may fit historical periods well even when those periods do not represent the test period.
  • Regime change is a possible explanation, not a confirmed diagnosis for a specific dataset.
  • The response does not offer a procedure for detecting regimes or correcting the generalization failure.

Tags

Full text
# Cannot achieve generalization of machine learning model


# Cannot achieve generalization of machine learning model












I'm working on a balanced, binary classification problem in a time-series (financial) dataset. I am using K-fold cross validation that is adapted for time-series (so that I'm never using future data to predict past data).

I have tried many algorithms, such as SVM, RandomForest and K-Nearest Neighbors. While all of them can achieve good results in cross validation, NONE of them have generalized well to the test set.

I use the cross validation to run grid-search feature selection and hyperparameter tuning simultaneously to find the best combination, but again - I have not achieved any generalization.

Do you have any ideas as to why this might be? Any general advice for dealing with this kind of scenario?

## Answer by demully (score 2)

https://quant.stackexchange.com/a/60831

One obvious answer, but not the one you'll probably want to hear is that maybe there is no single distribution for the variable you're trying to explain. It might be for want of a less-bad metaphor "regime-driven".

Being simplistic for simplicity's sake, imagine you were looking at the relationships between stock, bond, and commodity markets. The period between ~2002-~2007/8 might be described as "China-on, China-off"; that 07/08 and 12/14 as "risk-on, risk-off"; 12/14 and 18/20 as "liquidity-on, liquidity-off"; and who knows to nature of this now? :-)

It would then be very possible to train any model for any historical sample, and achieve attractively comfortable cross-validation results in your training set. However, generalisation in your test set might nevertheless stink, if the test set represented a different regime, with a different set of prevailing norms, assumptions, and/or trading rules.

I've certainly seen this in many of my own models over the last decade.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.