Skip to content
All library documents

Diagnosing Machine Learning Models That Fail on New Market Data

Article Quant Q&A · Author: dgmattam

Summary

The document describes a trading-modeling question in which neural networks and support vector machines show strong internal validation accuracy on three-class labels but perform near chance on a separate dataset. The models use many technical and fundamental inputs, raising concerns about whether the reported validation results reflect genuine predictive power.

Responses identify several possible causes: a change in the data distribution, overfitting, excessive or correlated features, high dimensionality, and inadequate normalization or preprocessing. Suggested approaches include decorrelating inputs, representing lagged price data with wavelets, adding summary statistics, and testing information from related markets. These are proposals rather than demonstrated fixes; one respondent reports similarly poor outcomes. The discussion gives no controlled comparisons or trading-performance evidence, and it does not establish which explanation applies to the original model. It underscores the need to scrutinize data splits and out-of-sample generalization in financial prediction.

Key ideas

  • High accuracy on carved-out samples can coexist with near-chance performance on later or separate market data.
  • Distribution shifts, overfitting, correlated inputs, dimensionality, and preprocessing may all contribute to poor generalization.
  • Decorrelated or transformed features, including lagged wavelet inputs and information from related markets, are proposed as possible alternatives.
  • The suggested remedies are not validated by comparative results in the discussion.

Tags

Full text
# Machine Learning on matlab 2010


# Machine Learning on matlab 2010












I am trying to develop a trading model. It uses certain technical and fundamental features and the model learns from the past. I have a 3-class output - bullish, neutral and bearish.

On trying neural networks, I got a train accuracy of around 85%, cross validation accuracy of 75% and test accuracy of 75% with both CV set and test set carved out from the complete set. I am doing it on matlab 2010 using nprtool. The software does the carve-outs of the CV set and test set (20% each and the rest of 60% is used for training).

After training I use the same model to test on a different data. Here the accuracy drops to around 34% (which I would assume is just random classification with 3 classes). The new data I test with is very similar to the data used to train. I also tried swapping the data sets, but the results are same. I used the command below to test with the model (the same was used to test the accuracy of the train data too with Xtest replaced with Xtrain). Here net denotes the trained model.

`output = sim(net, Xtest');`

For NN I used 100,000 samples, 250 input features and 100 nodes in the hidden layer.

I tried a similar thing using libsvm (called from matlab). Here too I face the same. During training, the parameters are optimized through cross validation and the CV accuracy comes up to 73%. However when tested on new data the accuracy drops to around 35%.

For SVM I trained with 20,000 samples and 250 features. I use the command below:

`[outputTest, accuracyTest, prob] = svmpredict(Ytest,Xtest,model);`

Please help if someone has come across a similar issue and has a remedy.

## Answer by user6430 (score 0, accepted)

https://quant.stackexchange.com/a/9807

Firstly, look at Jurik Research WAV and DDR modules to see one particular approach for time series data compression and decorrelation prior to input into ANNs. From recent research, it also seems like better results have been obtained using wavelets on lagged (2,5,10 bar) data along with indicators and summary statistics such as skewness, kurtosis.

You have too many inputs, and if there is any correlation between them, you will be defeating the purpose of ANNs. A basic rule of ANNs is that they will waste time learning the correlation between input features, so if it is removed and orthogonal inputs (zero correlation) are used, the results may be better.

Last, intermarket analyses has also produced better results than staying within assets. That is, start looking at other indices (lagged and unlagged) as inputs.

Don't fret over poor results -- I have obtained similar results using normalized data, 4 wavelets representing highs and four wavelets representing lows (and a variety of lags) for a 2-class problem (tomorrow is a gain(y>0) or loss(y<0)) with very poor results.

## Answer by joshi (score 0)

https://quant.stackexchange.com/a/9806

There can be several reasons for this:

- The `"new data"` that you use post-training & post-validation is `not drawn` from the same distribution as the one that you used to create/draw your training, testing and validation data.

- Since you have not mentioned anything related to the input features in your data-set, I am assuming that the `stock/option/derivative/instrument price` is one of them. I also assume that you know that the price movement is `stochastic` in nature (or at least the current accepted model is that price movements exhibit a stochastic random-walk behavior). You need to account for this. You can take a look at any standard textbook on quantitative finance (e.g. Wilmott).

- You might be `overfitting the model`

- The `curse of dimensionality`.

- Your data is not `normalized/preprocessed` in the right way. Again, I can't say more as I don't know how your data looks like.

- In principle, it is a very HARD problem. If it was easy, well, all of us'd be retired. I mean there is `no free-lunch`. :)

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.