Methods for Selecting Factors to Explain Mutual Fund Returns
Summary
The document considers how to identify which candidate factor returns help explain mutual fund returns. It presents factor selection as a feature-selection problem and lists several approaches: rank factors using individual correlations, use recursive feature elimination with cross-validation, or apply backward or sequential elimination. Random forests are suggested because they can provide feature-importance measures; neural networks are mentioned as another possible predictive model. Tools such as SHAP can help interpret model contributions.
Other responses suggest examining correlation or R-squared, and applying principal component analysis to a group of fund returns before comparing the resulting components with predefined factors. The discussion also favors decomposing the returns matrix using a chosen set of predefined factors when interpretability matters. These are suggestions rather than a worked comparison: no fund data, validation results, or selection criteria are supplied. Predictive feature importance and statistical association do not by themselves establish that a factor is economically causal or stable out of sample.
Key ideas
- Factor selection can be framed as supervised prediction of fund returns from factor returns.
- Correlation ranking, recursive elimination, and backward selection are candidate feature-selection methods.
- Random forests can provide feature-importance rankings, while interpretability tools can help examine model contributions.
- PCA can summarize variation across funds, but its components may not map directly to predefined economic factors.
- The document gives methodological suggestions without comparative results or evidence of out-of-sample stability.
Tags
Full text
# Factor selection for predicting fund returns # Factor selection for predicting fund returns I have a list of factors (and their returns) as well as a set of mutual fund returns. What are some techniques I could use to select relevant factors for the funds. For example, fixed income factors to be selected for fixed income funds. I have tried stepwise regressions and then filtering on the $p$-value, but I was wondering if there are other methodologies. Would neural nets be a candidate for this? ## Answer by alexprice (score 2) https://quant.stackexchange.com/a/54702 You can formulate this as a machine learning problem of predicting mutual funds return based on factor returns. Any machine learning model can be used, such as neural nets, although tree based models such as Random Forest would be more suitable as they provide feature importance. Then your problem of selecting relevant factors is known as "feature selection" for which you can start by using univariate methods such as described in https://scikit-learn.org/stable/modules/feature_selection.html (i.e. calculating correlation and rank features by correlation) , or more time-consuming solutions such as RFECV or backward feature elimination such as http://rasbt.github.io/mlxtend/user_guide/feature_selection/SequentialFeatureSelector/ . Also you can look into eli5 and SHAP libraries. ## Answer by Gogo78 (score 0) https://quant.stackexchange.com/a/48663 you can also use correlation,R2 to detect pairs. ## Answer by Dhruv Mahajan (score 0) https://quant.stackexchange.com/a/49180 I'm assuming you want to check factors for a group of funds (from what you've written in the question) and not a single fund. Simple thing to do would be PCA on the set of different mutual fund returns and get the factors explaining most variance. Again PCA factors are completely black-box so now you'll have to check correlation of the highest variance explaining factor (PC1) with the group of pre-defined factors you want to check for. Better approach is to decompose the returns matrix based on the set of pre-defined factors. There is a paper by Meucci in it. If you're good at matrix algebra give it a read : Paper
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.