Alternative Data and Machine Learning for Investment Research
Summary
This post summarizes a 2017 J.P. Morgan report on using large and alternative datasets in investment research. It groups alternative data into information generated by individuals, business processes, and sensors, with examples such as social media, commercial transactions, and satellite imagery. The report’s framework covers acquiring and interpreting data, then applying statistical or machine-learning methods. These include supervised regression and classification, unsupervised clustering and factor analysis, and deep or reinforcement learning. It also surveys data and technology providers.
The article argues that alternative data can provide more timely signals than traditional releases, but turning raw data into trades requires substantial organization, analysis, and market understanding. It warns that datasets may be costly, weak, or quickly lose predictive value; complex models can overfit or find spurious patterns; and technical sophistication alone does not establish an economic rationale. The summary is based on an older report, notes limited deep-learning coverage and few financial applications, and offers no systematic evidence of investment performance.
Key ideas
- Alternative data can come from individuals, business processes, and sensor systems.
- The report surveys supervised, unsupervised, deep-learning, and reinforcement-learning methods for analyzing such data.
- Data acquisition and interpretation are central parts of the investment process, alongside model selection.
- A predictive relationship must be evaluated for economic meaning and converted into a tradable strategy.
- Data costs, signal decay, overfitting, and spurious patterns limit the approach.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.