Skip to content
All library documents

Choosing Predictors: Multicollinearity and Out-of-Sample Validation

Article Quant Q&A · Author: Tal Fishman

Summary

The discussion addresses how to decide whether a predictive model has too many explanatory variables, especially when adding variables keeps improving in-sample fit. The answers emphasize that there is no universal cutoff. One warning sign is multicollinearity: if adding or removing a regressor substantially changes other coefficients, the estimates may be unstable. A model with lower explanatory power can be preferable when its coefficients are more reliable.

Suggested checks include reserving holdout data and evaluating out-of-sample performance, while keeping that evaluation for late in the research process because repeated experimentation can overfit the holdout too. Adjusted R-squared is also mentioned as a way to account for the number of variables, though it is not presented as a definitive test. The discussion offers practical guidance rather than a formal selection procedure or trading results. Its central limitation is that model complexity cannot be reduced to a single metric; validation depends on disciplined research and the question being tested.

Key ideas

  • Adding predictors can improve in-sample fit without improving live performance.
  • Large coefficient changes when regressors are added or removed can signal multicollinearity.
  • Unstable coefficients may make a lower-fit model more useful than a model with higher R-squared.
  • Holdout data can assess out-of-sample performance, but repeated tuning can overfit it.
  • Adjusted R-squared accounts for the number of predictors but does not provide a universal cutoff.

Tags

Full text
# How many explanatory variables is too many?


# How many explanatory variables is too many?












When researching any sort of predictive model, whether using ordinary linear regression or more sophisticated methods such as neural networks or classification and regression trees, there seems to always be a temptation to add in more explanatory variables/factors. The in-sample performance of the model always improves, and sometimes it improves a great deal, even after one has already added quite a few variables already. When is it too much? When is the supposed improvement in in-sample performance very unlikely to carry over into live trading? How can you measure this (beyond simple things like the Akaike and Bayesian Information Criteria, which don't work very well in my experience anyway)? Advice, references, and experiences would all be welcome.

## Answer by Richard Herron (score 4, accepted)

https://quant.stackexchange.com/a/1681

“Make things as simple as possible, but not simpler.” The problem you want to avoid is (near) multicollinearity. The tip-off will be that adding/removing a regressor will significantly change the coefficients on the other regressors. In practice (well, in the research that I read) I rarely see this explicitly tested.

If you think that you have multicollinearity, then it's likely best to either estimate over a subset without multicollinearity or to drop the offending regressors. A model with less explanatory power as measured by $R^2$ is certainly better than a model with incorrect (unstable) explanatory power.

## Answer by wburzyns (score 4)

https://quant.stackexchange.com/a/1678

Although not directly related to financial modeling, I've found the following quotation to be very instructive:

"I remember my friend Johnny von Neumann used to say, 'with four parameters I can fit an elephant and with five I can make him wiggle his trunk.'" -- E. Fermi

You may also read this: http://mahalanobis.twoday.net/stories/264091/

## Answer by Ari B. Friedman (score 3)

https://quant.stackexchange.com/a/1680

There's no rule to answer this question for you. You need some combination of:





- Hold-outs: You correctly mention that the problem is "in sample performance." The solution is therefore to hold out some data when you start and look at out-of-sample performance. Of course, if you iterate enough times, you can over-fit your holdout sample, too! So save this until the last step, and be honest with yourself.

As always, the key is to be certain of what question you are trying to answer. Then you can muster as much unbiased evidence as possible.

## Answer by WaveRider (score 2)

https://quant.stackexchange.com/a/1823

I think you're looking for a metric that quantifies the effectiveness of the added variable(s). Objectively, you want each variable to have correlation to your model estimation output and non-correlation between other variables that may be utilized. If you adjust your $R^2$ metric accordingly (less degrees of freedom per variable) you'll get a reasonable feel for where the limit is for adding more variables (otherwise $R^2$ will just increase and you're back to asking the same question).

## Answer by Michael (score -2)

https://quant.stackexchange.com/a/1786

Just pick up a decent econometrics book (Gujurati is what I used in school).

If you have multicollinearity, find a dummy variable.

http://en.wikipedia.org/wiki/Coefficient_of_determination#Adjusted_R2 << this should be somewhat helpful.

I have no trading experience, so cum grano salis.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.