Scaling Price Features Across Stocks for Machine Learning
Summary
The document discusses how to make price-derived features comparable across stocks with different price levels and across periods when an instrument’s price changes substantially. One answer recommends log returns as a common scale, followed by standardization such as centering and scaling by standard deviation. It cautions that return magnitude may itself be predictive, so normalization can remove information a supervised model should be allowed to learn.
A second suggestion is to add a volatility-adjusted price feature by dividing price by an estimate of volatility, calculated from a trailing historical window or obtained from implied volatility. The response recommends retaining the unadjusted price as a separate feature because the adjusted value does not uniquely identify the original price. These are brief suggestions rather than a tested comparison: the document provides no dataset, model results, choice of lookback window, or guidance on preventing look-ahead when estimating volatility.
Key ideas
- Log returns can put price changes from differently priced stocks on a more comparable scale.
- Standardizing returns may also remove magnitude information that could be useful to a model.
- A volatility-adjusted price feature can divide price by a historical or implied volatility estimate.
- Keeping original price alongside an adjusted feature preserves information that the ratio may omit.
- Historical volatility features should use estimates available at the time of each observation.
Tags
Full text
# Right way to standardize price based features across different stocks for supervised learning # Right way to standardize price based features across different stocks for supervised learning Let's say we have an OHLCV dataset for a universe of stocks. We want to create features based on these price data. Since each stock may have a very different price range from the other if we just take log-delta (eg. open-close) a stock that has a price range around 1000-2000 would look very different from another that has a price range of around 1-10. This will happen also for stock (or any instrument) that has a drastic shift in prices over time. For instance, BTC recent steep rise in value. What would be a good way to standardize or normalize these features across different stocks of various price ranges? ## Answer by spark (score 2) https://quant.stackexchange.com/a/66147 Log returns are usually sufficient to place different stocks on the same scale. Yes, some stocks may have 100% returns when others have 1%. And you can standardize them (mean=0, stdev=1). But, isn't this a feature that you want your model to capture rather than remove by normalization or standardization? ## Answer by Ralph Winters (score 0) https://quant.stackexchange.com/a/70615 @hugh. You might try to add a feature by standardizing the stock price via dividing the price by the standard deviation, similar to how a Sharpe ratio is used to factor in volatility when comparing historical returns. When looking at the price data historically, you would need to get an estimate of the volatility which existed for each price data point. You could do this by calculating the standard deviation for the historical data point over the previous x number of days, or use an implied volatility. I would also keep the original stock price in the model since price/standard deviation not unique for each stock price, but adds value as a feature.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.