Selecting and Cleaning Historical Stock Data for Machine Learning
Summary
The answer offers practical checks for choosing equity data used in technical-analysis and machine-learning research. It advises excluding thinly traded stocks when recorded prices may not reflect executable prices, recognizing duplicate economic exposure from ADRs and local listings, and accounting for stocks whose behavior changes after a fixed-price acquisition announcement. It also suggests checking collinearity when selecting related securities.
Data timing matters when combining markets: closing or settlement prices can refer to different clock times, so a strategy may mistakenly assume it could trade on information that was not yet available. Historical datasets should also retain delisted or acquired companies that were actually listed at the time being studied. These points help reduce tradability errors, look-ahead problems, and survivorship bias. The answer gives screening principles rather than a complete dataset-selection procedure, and it does not prescribe a universal universe or quantify the impact of each filter.
Key ideas
- Thinly traded stocks can have recorded prices that are difficult to trade at in practice.
- Duplicate listings can create misleading appearances of distinct securities or pairs.
- Acquisition announcements can materially change a stock’s behavior and may justify removing it from analysis.
- Price timestamps must be synchronized to avoid using market information before it was available.
- Historical datasets should include securities that later delisted or were acquired to limit survivorship bias.
Tags
Full text
# Answer by JoshK (score 4, accepted) # Should there be a relation between stocks when used as input data for integrating Technical Analysis with Machine Learning? I'm integrating Technical Analysis with Deep Learning for the first phase of my research. I wanted to know how should I pick (or group) stocks as input data and whether there should be relation between the selected stocks. To further elaborate, I've seen researchers use different stocks, some eliminate company stocks below certain market cap, others use the whole historical price chart of S&P 500, and I can't find the reason behind their choice. Is there a best practice for selecting the data sets or should I just do it intuitively? ## Answer by JoshK (score 4, accepted) https://quant.stackexchange.com/a/42410 There are a few exclusions that I have commonly seen: - Excluding thinly traded stocks. The price that shows up in your data feed may not relate to actual tradable prices. - Filtering for ADR/Pink locals. You can find stocks listed in multiple places in ways that would lead you to think that they are great for pairs trades when actually they are the same stock, but just with listing differences. For example CS (Credit Suisse NYSE ADR) and CSGKF (Credit Suisse Pink Sheet Local). Screening for co-linearity can be helpful as well... - Removing stock post corporate acquisition announcement. Once a stock is being acquired for a fixed amount it will lose many of the properties that you are trying to analyze. - Handle time synchronization issues. Some data sets will show you "close/settlement prices" that are taken at different points in time. For example, a US equity close price is taken at 4:00pm EST, an oil contract close price is taken at 2:30pm EST, and a bond future close price is taken at 3:00pm. If your algorithm tells you to buy oil if XOM closes above it's 20 day moving average and it thinks that you could have transacted at 4:00pm at the 2:30 price, then you can imagine the errors that will occur. And one important thing to screen back in: Many data sets will drop delisted / acquired stocks. When analyzing historical data you need to make sure that your data set includes your candidate names that were actually trading at the time.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.