Skip to content
All library documents

Relating Model Complexity to Sample Size and Overfitting

Article Quant Q&A · Author: Gabriel Gómez Rojo

Summary

The document asks whether there is a general way to determine how much data a model needs relative to its number of free dimensions. The author is concerned that adding variables may improve apparent accuracy while also increasing the risk of curve fitting. A proposed heuristic is to use at least thirty observations per variable, but the author is unsure whether that ratio is justified and notes that model type and relationships among predictors may matter.

The responses point to information criteria such as Akaike’s Information Criterion for judging fit quality, and to Vapnik–Chervonenkis theory as a broad theoretical framework for learning and generalization. They also caution that such theory is difficult and that practical heuristics may be more useful. The document offers no universal ratio, worked analysis, or trading-strategy validation evidence; sample requirements remain dependent on the model, data structure, and validation objective.

Key ideas

  • More model variables can improve apparent fit while raising the risk of overfitting.
  • A fixed observations-per-variable ratio is presented as an uncertain heuristic, not a general rule.
  • Akaike’s Information Criterion is suggested for comparing fit quality while accounting for complexity.
  • Vapnik–Chervonenkis theory offers a broad theoretical lens on model capacity and generalization.
  • Model class and predictor relationships affect how much data may be needed.

Tags

Full text
# How does the number of free dimensions of a model affect its required size of sample?


# How does the number of free dimensions of a model affect its required size of sample?












Adding more variables to a model usually increases its accuracy. However, without adequate analysis it could also lead to curve fitting.

Another question (How much data is needed to validate a short-horizon trading strategy?) received answers related to the statistical significance of the standard error of the model. However, I wonder if anyone has results (or analysis) of what should be the ratio of sample data to dimensions used in a model. My intuition has led me to use at least 30 times more sample data points than variables implemented as dimensions but I am not happy with this approach.

I guess that this would depend on the characteristics of the model (it would be different for linear regressions, SVM, non-linear models, etc. and also dependent on the relationships among the variables used) but is there a general framework for estimating this?

## Answer by Andrew (score 3)

https://quant.stackexchange.com/a/7615

The following is a good way to judge the quality of fits for a model.

http://en.wikipedia.org/wiki/Akaike_information_criterion

## Answer by g g (score 3)

https://quant.stackexchange.com/a/7651

In full generality this is a very difficult question. The closest you will get to a general framework is Vapnik-Chervonenkis theory. You can read about this in Chapter 7.9 of "The elements of statistical learning" by Hastie, Tibshirani and Friedman which can be downloaded from their website .

But be warned that this is a theoretical approach. Often more heuristic approaches will serve you better. Chapter 7 of the book covers those as well.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.