Skip to content
All library documents

Genetic Programming for Nonlinear Equity Factor Discovery

Article BigQuant

Summary

This research summary outlines three refinements to genetic programming for finding stock-selection factors: fitness measures based on mutual information and long-only excess return, ways to transform nonlinear factors, and validation procedures intended to reduce overfitting. Mutual information can capture nonlinear links between factor values and returns. The reported factor tests often show higher returns in middle-ranked groups than at either extreme, a pattern that machine-learning models may use. The work also reports more than twenty discovered factors, though the supplied text does not provide detailed performance statistics.

For nonlinear factors, it contrasts fitting their relationship to returns directly with machine learning against transforming each factor for a linear model. It describes cubic-regression residual and polynomial-fitting approaches, with the latter reported to work better but requiring factor-specific fits. Its validation workflow evolves factors on a training set, measures each generation on a validation set, and stops when average validation fitness converges. The evidence is limited to Chinese A-share stocks; discovered patterns may fail, and complex expressions can reduce interpretability or generalize poorly beyond the tested universe.

Key ideas

  • Mutual information can guide genetic programming toward factors with nonlinear relationships to returns.
  • A long-only excess-return objective can discover factors optimized for long positions.
  • Cubic-residual and polynomial transformations offer different trade-offs for linear factor use.
  • Validation-set fitness across generations provides a way to monitor convergence and overfitting.
  • The reported evidence concerns A-shares and does not establish performance in other markets.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.