Skip to content
All library documents

Selecting Factors by Predictive Power and Investment Performance

Article Quant Q&A · Author: Ram Ahluwalia

Summary

The discussion surveys ways to rank candidate variables for factor models when correlation among inputs is not the main concern. Suggested measures include rank correlation, causal or cointegration tests, information ratios, regression fit, premium persistence and volatility, monotonicity across sorted portfolios, and the significance and economic performance of long-short spreads. Replies add variance reduction, hit rate, turnover, lagged correlation decay, dimensionality-reduction methods, information criteria, and predictive checks.

The answers offer competing views rather than a settled selection procedure. One contributor favors the factor’s t-statistic and gives a rule of thumb, while another describes principal components regression and partial least squares as modeling approaches. The discussion does not provide a common dataset or empirical comparison showing that any single metric dominates. In practice, the suitable measure depends on the target, model assumptions, trading costs, and whether the goal is explanation or investable performance; screening results also need validation to limit overfitting.

Key ideas

  • Factor candidates can be assessed with statistical association, explanatory fit, and portfolio performance measures.
  • Sorted portfolio spreads can be examined for significance, monotonicity, hit rate, and persistence.
  • Turnover and decay in predictive relationships help assess factor stability after trading costs.
  • Principal components regression and partial least squares offer ways to build models from many inputs.
  • The responses disagree on whether a t-statistic or a broader set of criteria should guide selection.

Tags

Full text
# Variable Selection in factor models


# Variable Selection in factor models












Let's say you have a dependent variable and many independent variables. What are the preferred metrics for sorting and selecting variables based on explanatory power? Let's say you are not concerned with correlation among your inputs.

My thoughts are below. Let me know if I'm missing anything obvious, or if there is a metric that you feel dominates the others.

- Spearman correlation controlling for sector

- Granger causality

- Johansen test for cointegration

- Grinold's Information Ratio

- Economic performance of the factor during a factor backtest (i.e. drawdown, sharpe, etc.)

- Volatility of the factor premium

- Persistency of the factor premium

- Monotonic relation test

- Significance test on spread return (Q5-Q1) vs. Mean Return

- R^2 of factor after building some kind of model

My preference would be correlation coefficient.

## Answer by I-CJW (score 4, accepted)

https://quant.stackexchange.com/a/1370

I'd add:

- Variance reduction

- Fraction same sign / Hit rate

Additionally, you might look at the relationship between the Q5-Q1 spread itself and the dependent (i.e. are larger/smaller spreads associated with some feature of the dependent).

Turnover may also be an issue as slippage and friction come into consideration. Measures such as percent turnover in the Q5-Q1 portfolio, and correlation coefficient decay over lagged periods can prove insightful in selecting factors with higher stability and persistence.

## Answer by bill_080 (score 1)

https://quant.stackexchange.com/a/1373

If you're using R, you might try:

http://cran.r-project.org/web/packages/relaimpo/index.html

https://stats.stackexchange.com/questions/8918/is-there-a-way-to-optimize-regression-according-to-a-specific-criterion/8932#8932

## Answer by pyCthon (score 1)

https://quant.stackexchange.com/a/16648

I don't think any of the current answers, answer the question, there is one metric that dominates all of them and that is simply the t-stat of the factor. This is well shown across academic literature as well.

The rule of the thumb here is to use a t-stat greater than 2.

## Answer by John (score 0)

https://quant.stackexchange.com/a/16683

There are a few other techniques to consider in addition to what has already been suggested.

Principal components regression (PCR). This would consist of doing a PCA on the independent variables and then regress the dependent variable on the factors. You can always translate the coefficients back into the original basis.

Partial Least Squares (PLS). Very similar to PCR, but while the PCR PCA is about choosing coefficients to maximize the explaned variance of the independent variables, in PLS you choose coefficients so that the independent variables explain the most variance of the dependent variable.

Information Criterion. If you can express the choice as between a simpler model versus a more complex model, the optimal choice under this approach is to choose the one with the most favorable information criterion. Most common is AIC, which increases in log likelihood of the model and decreases in number of paramters. Of course, this requires making some assumptions about the errors of the model.

Posterior predictive checks. Simulate data from the model. Evaluate if it makes sense.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.