Skip to content
All library documents

Checking Model Overfitting with Train-Test IC Consistency

Article BigQuant

Summary

This article proposes comparing a model's information coefficient (IC) and mean IC across training and test samples to assess predictive stability. It defines IC as the correlation between a model score and subsequent returns, using a five-day return example, and treats the sign as an indication of predicted direction. A model with similar sample characteristics and consistent IC direction across training and test periods is presented as evidence of better generalization.

The author emphasizes fixing the test period in advance and avoiding repeated adjustments to make its backtest look better. A near-zero IC or weak training signal may indicate limited predictive power, while training results alone cannot reliably identify overfitting. The piece illustrates its approach with an example where both samples show negative IC, but does not provide independently verifiable statistics or a formal threshold. It also argues that live performance matters more than an attractive historical curve; the examples and claims about future results are anecdotal, so IC consistency is a diagnostic rather than a guarantee.

Key ideas

  • Compare training and test IC values and signs to assess whether a model's predictive direction persists out of sample.
  • Define the test period in advance and avoid tuning the model repeatedly against it.
  • IC measures the relationship between model scores and subsequent returns, with its sign indicating direction.
  • A near-zero IC can signal weak predictive content, while training data alone cannot establish overfitting.
  • Consistent IC direction is presented as a diagnostic, not proof of future profitability.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.