Skip to content
All library documents

Data Preparation and Testing Choices for Equity Multi-Factor Models

Article BigQuant

Summary

This summary of a brokerage research report discusses the research workflow for multi-factor equity selection. It highlights source-data issues such as delayed or unreliable financial disclosures, corporate restructurings that disrupt historical comparability, and incomplete industry classifications. It also describes shaping the stock universe to suit the strategy, including a liquidity-based universe as an alternative to using the full sample.

For cross-sectional preprocessing, the document recommends handling outliers before imputing missing values and standardizing factors, since these later steps may use cross-sectional averages. It contrasts z-score scaling, which retains information about distances between factor values but remains sensitive to extremes, with rank scaling, which is more robust to extremes but discards those distances. It identifies long-short portfolio sorts and regression as common single-factor predictive tests. The summary warns that findings based on models and historical data may fail to persist; it provides no detailed results from the underlying report.

Key ideas

  • Financial data can be delayed, unreliable, or hard to compare across corporate restructurings.
  • The stock universe should be selected to fit the intended strategy, and liquidity can guide its construction.
  • Outlier treatment should precede missing-value handling and factor standardization.
  • Z-score scaling preserves cross-sectional distances but is sensitive to extreme values, while rank scaling discards distances and is more robust.
  • Long-short sorts and regressions are presented as ways to assess a factor’s ability to predict returns.
  • Historical model results may not persist in future markets.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.