Skip to content
All library documents

Comparing Normalization Methods for Market Returns

Article Quant Q&A · Author: ninjaSurfer

Summary

This document poses a factor-modeling question: how to put returns measured on different scales onto comparable scales before predicting target returns. It lists several candidate methods: subtracting the mean and dividing by standard deviation, dividing by standard deviation alone, scaling by the Euclidean norm, and using the dispersion of highs and lows.

The text does not compare these methods, report empirical tests, or recommend one. It highlights a practical concern in noisy market time series: centering assumes drift is meaningful, while different scale estimates may respond differently to noise. The document is therefore a research question rather than evidence for a preferred normalization. Any choice would need to be assessed in the intended factor model and on suitable out-of-sample data.

Key ideas

  • Factor returns on different scales may need normalization before use in a target-return model.
  • Candidate approaches include standardizing around the mean, scaling by standard deviation, using Euclidean norm, or estimating scale from highs and lows.
  • Mean subtraction can be problematic when the assumed drift is not meaningful.
  • The document supplies no comparison, empirical evidence, or recommendation.

Tags

Full text
# Comparison of normalization methods on market returns


# Comparison of normalization methods on market returns












I am looking to use a multi-factor model to make target-return predictions. Since the factor-returns come from different scales I need to normalize first.

There are different ways to normalize returns, to mention a few: subtract mean and divide by standard deviation(assumes non-zero drift), simply divide by standard deviation, divide by euclidian norm, divide by the standard deviation of highs/lows.

My question: is there a documented comparison of the different methods, on noisy time-series data such as market returns, stating the pros/cons of different methods? Empirically do you have any suggestions and remarks?

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.