Curating Alternative Data for Quantitative Trading
Summary
The document surveys data sources that systematic traders can use beyond historical prices, including traded volume, company fundamentals, and alternative data such as satellite imagery, news text, supply-chain information, and card transactions. Its main lesson is that adding data does not automatically improve a trading model: alternative datasets can contain biases introduced when observations are mapped to companies or other economic entities.
Examples include linking a satellite image to a business activity, matching shop names to companies, or assigning news text to an issuer. These mappings can miss entities or misidentify them, particularly when brands, inventory ownership, or company relationships change. The document recommends identifying data issues close to collection, using domain expertise and descriptive analysis, and preserving point-in-time mappings so historical research reflects what was knowable at the time. Machine-learning classifiers may assist with entity matching, but the source emphasizes that they do not remove the need to understand the dataset. It offers principles and examples, not a quantified comparison of data sources or evidence that any one source creates an edge.
Key ideas
- Systematic trading can draw on prices, volume, fundamentals, and alternative datasets.
- Entity matching can introduce systematic errors into alternative data.
- Mappings should reflect information available at each historical point to reduce survivorship bias.
- Domain expertise and descriptive analysis are important parts of data curation.
- Machine-learning classifiers can assist matching but do not replace dataset expertise.
Tags
Full text
# What kinds of data should be curated for day trading? # What kinds of data should be curated for day trading? From previous research about data curation with research papers, it seems to me that most algorithmic trading systems (at least in regards to day trading) solely use historical price data- but I'd be interested to see if there were other sources that could help gain an edge. When it comes to day trading, do we just need to solely input *historical price data for our models or could other sources help as well? If so, what are these sources? And would they help more than hinder? [Historical Price Data Example: Open, High, Low, Close] Thank you! ## Answer by lehalle (score 3) https://quant.stackexchange.com/a/69187 Historically systematic trading strategies have used - historical price data (and traded volumes) - fundamental data (balance sheets of companies) - alternative data (satellite images, texts, supply chain, credit cards, etc) the last 10 years have seen the emergence of alternative data, few references - The Book of Alternative Data: A Guide for Investors, Traders and Risk Managers, by A Denev and S Amen; - Big Data and Machine Learning in Quantitative Investment, by Tony Guida. One important point with alternative data is that they have a lot of biases. The role of Data Curation is to identify these biases and to correct them (when it is possible). It involves descriptive statistic, expertise of specific fields, and data modelling. The are two kind of biases (see my talk "Biases of Learning Machines in Finance: Some Examples" at the 2nd ACM International Conference on AI in Finance) - Observed entities, due to the "matching" Satellite images and geolocation: a polygon of Lat x Lon gives you an Area of Interest, that you have to match with an economic activity shares by economic entiries (like “this is a field of corn” or “this is a Starbucks”) Credit cards: you need to match shop names and brands to companies, Financial News: you need to match a paragraph of a News to a company or an economic entity, Etc This matching introduces biases: you can systematically miss Companies who are not owning their inventories but lending them A misspelled brand name, or miss-match a brand that has been recently sold to another company Etc And all this matching has to be Point in Time to be able to replay the past without any survivorship bias. This mapping can be enhanced using machine learning algorithms (f.i. classifiers). These biases should be identify as close as possible to the data collection, it cannot be done without a field expertise in domain documented by the dataset.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.