Skip to content
All library documents

Selecting Active Mutual Funds with Dynamic Machine Learning

Article BigQuant

Summary

This research review describes a dynamic, out-of-sample process for selecting US active equity mutual funds from observable fund characteristics. It compares elastic net, random forests, gradient boosting, and linear regression. Each year, models are trained on prior data to forecast fund alpha; an equal-weighted portfolio is formed from the predicted top decile and evaluated against factor models and passive benchmarks. The dataset covers fund share classes from 1980 to 2018, and the review reports positive, statistically significant risk-adjusted performance for gradient boosting and random forests, while elastic net and ordinary least squares do not achieve significant positive alpha.

The reported advantage depends on nonlinear relationships and interactions across multiple predictors, whose importance changes over time. Robustness checks vary portfolio cutoffs and performance models and include retail-only funds and neural networks. These findings are historical and do not guarantee future returns: all portfolios’ alpha declines toward the end of the sample, and the text notes that predictability may disappear as markets become more competitive. The reported results are also specific to US mutual funds and the study’s data and modeling choices.

Key ideas

  • The study compares linear and machine-learning models for forecasting active mutual fund performance from lagged fund characteristics.
  • Gradient boosting and random forests produce significant positive risk-adjusted results in the historical out-of-sample exercise; elastic net and OLS do not.
  • Using several predictors and modeling nonlinearities and interactions works better than relying on only a few features.
  • Predictor importance changes over time, supporting periodic model refitting.
  • Alpha declines late in the sample, so historical predictability may weaken or vanish.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.