Skip to content
All library documents

Scaling Technical Indicators for Cross-Asset Machine Learning

Article Quant Q&A · Author: siamii

Summary

The discussion considers how to scale a large set of technical indicators when training a model across many stocks with different price levels and feature ranges. Suggested approaches include range scaling, mean-zero standardization, and converting skewed feature values to ranks, percentiles, or normal scores. It also notes that the best choice depends on the model: tree ensembles may be less sensitive to scaling than recurrent neural networks.

No comparative experiment is reported. One answer recommends selecting preprocessing by out-of-sample profit and risk, while another describes dividing recurrent-network inputs by a sample-wide standard deviation, conditional on the inputs being approximately normal. The advice is exploratory rather than a definitive recipe: the thread does not specify leakage-safe fitting procedures, a common evaluation design, or how indicator transformations interact with time-series and cross-sectional structure. The original question also raises whether indicators should be calculated before or after normalizing prices, but the replies do not resolve that point.

Key ideas

  • Indicator ranges can differ substantially, so scaling may prevent large-valued features from dominating some learning methods.
  • Range scaling and standardization are common options, but skewed data may call for rank or percentile transformations.
  • Input scaling sensitivity varies by model, with recurrent networks described as more sensitive than random forests.
  • Choose preprocessing using out-of-sample performance that accounts for both profit and risk.
  • The discussion does not establish a universally correct transformation or settle whether to normalize prices before calculating indicators.

Tags

Full text
# How to normalize technical indicators for machine learning?


# How to normalize technical indicators for machine learning?












I'm using around 130 technical indicators for 100 different companies. Each company's stock price moves in a different range, see FTSE 100. In addition, each technical indicator moves in a different range as well, ie some goes between 0-1, other 0-100 and others move with the price. This is the list I'm using http://ta-lib.org/function.html

I'd like to feed this into a machine learning algorithm where I predict the relative price movement of the stock price the next day. I use a logistic loss for profit optimization and two regularizer terms, one between companies, and the other between time periods. This is unimportant for now.

Rather, what I'm asking about is how to normalize the input data? I've tried zscore, rate of change, absolute differences, and various combination of these, but I'm not sure which is the right approach. Also, I assume that first I need to calculate the indicators and then normalize that data. or would these indicators make sense on normalized data?

## Answer by user6430 (score 3)

https://quant.stackexchange.com/a/9839

Scale and range are your biggest issues. If one input has values which range from e.g. 2300-3500, and another from 0 to 18, then the large scale of the first will swamp the other and provide greater informativity into your learning algorithm. Therefore, normalize into range [0,1] or mean-zero standardize - like you have already done. Be careful with mean-zero standardization, however, since means only apply to skew-zero normal distributions, and not to log-normally distributed (right-tailed) skewed distributions. You could input the ranks of your skewed features values, but ranks are known to be rectangularly distributed, so convert ranks to percentiles, and then to van der Waerden scores (which are standard normal distributed without skewness).

## Answer by user2763361 (score 2)

https://quant.stackexchange.com/a/9848

There is no right approach a priori. Try all approaches that make decent sense and pick the one that maximises a utility function on out of sample PnL and risk (or some similar decision rule).

## Answer by wildbunny (score 2)

https://quant.stackexchange.com/a/44885

This depends on what ML you are using. For example, random forest is extremely resilient and doesn't require much normalisation of inputs to produce good outputs. However, RNN is extremely sensitive to input normalisation.

What I do for RNN is to divide each sample input by 1 standard deviation of the entire sample set, and that works quite nicely. But make sure your samples are normally distributed first.

## Answer by Jacques Joubert (score 1)

https://quant.stackexchange.com/a/44879

There are a couple of good papers on this but my favorite are Predicting stock market index using fusion of machine learning techniques and Predicting stock and stock price index movement using Trend Deterministic Data Preparation and machine learning techniques.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.