Skip to content
All library documents

Why Market Data Mining Produces Spurious Trading Signals

Article Quant Q&A · Author: davegaut

Summary

The document describes an attempt to find directional patterns in second-by-second market depth data using dimensionality-reduction and classification methods. A UMAP visualization appeared to separate observations labeled by subsequent price direction, but closer inspection and statistical analysis did not confirm a trend. The post asks how quantitative trading firms find profitable anomalies, particularly for market making.

The response argues for starting with a plausible model of how markets work and using data to challenge it, instead of searching raw data for patterns. Its example considers whether social-media mentions predict trading volume: both may reflect the number and behavior of shareholders, so mentions could add little information or create multicollinearity. The answer warns that accidental regularities can look meaningful. It offers general research guidance rather than a tested trading strategy, and the example does not quantify predictive performance or establish that any particular signal works.

Key ideas

  • A visually distinct embedding can create a misleading impression of predictable price direction.
  • Searching raw data for patterns can uncover many accidental correlations.
  • Begin with a market hypothesis and use data to try to disconfirm it.
  • Consider whether an apparent predictor simply reflects another underlying variable.
  • Correlated inputs can add little information and may cause multicollinearity.

Tags

Full text
# Insoluble Enigma


# Insoluble Enigma












I used many statistical tools, i.e. t-SNE, SVM, Neural Network, UMAP, PCA, with the reformatted full market depth data with timestamp each second. UMAP gave me the best data representation, but clearly, it was just an illusion.

Cloud 1

Cloud 2

The two above pictures showed me an illusion. I thought I might see two distinct clusters stick together. Be aware that each color is a pre-defined label. The green points are just label 1 meaning the price is going up and the red color are just the label -1 meaning the price is going down. There's also the label 0 meaning the price is stable, but I did not consider here

If I am going closer, I see the following picture

Cloud 3

I analysed the data statistically and I realized there was just no trend.

I got the most valuable data and I have never seen trend analytically. How edge fund firm succeeded to find anomalies in the data so that they can apply market marking and make profit from them?

## Answer by Dave Harris (score 6)

https://quant.stackexchange.com/a/43029

Your question is too broad to give anything but a very general answer. Data mining in the raw form won't do any good. At the minimum, you will pick up thousands of spurious correlations. You cannot go from data to a solution. You have to work in the opposite direction, you have to posit some model of the world and then test it. You must have an existing understanding of the relationships among the variables.

For example, let us imagine you believed there was a positive relationship between mentions on Twitter and volume traded. That may be true, but there are many lurking variables. Imagine two companies with equal capital bases and revenues, but with one difference. The first firm had one hundred distinct shareholders and the second firm had 5000 distinct shareholders. I say distinct because you want to ignore husband and wife, family office and related corporation combinations.

In the first company most shareholders, for whatever reason, are buy and hold shareholders, while the second trade quite a bit. If the second company had more comments on Twitter, well, that is unsurprising because they have more distinct shareholders. If the Twitter feed is a function of the number of stakeholders, then no new information about the trading pattern of the firm was added by using Twitter. Indeed, using Twitter may, in fact, reduce the information available as it is derived from the same dataset. It may be that counting shareholders is a superior metric than counting Twitter comments. Using both may trigger multicollinearity problems or trigger problems with the definition of the function. The Twitter comments may be a highly non-linear function of the count of shareholders.

You should never start in the data. You should start with your knowledge of the world, then use the data to disconfirm it.

Spurious relationships abound because accidental regularities abound. Think about the regularities in your life that look like they are not independent decisions, such as when you eat lunch. Notice that millions of people eat lunch at the same time as you. They do not have to and it isn't an intrinsic relationship. You could wake up an hour earlier, go to work before most people, and eat at lunch at ten in the morning. Those social accidents make it onto the tape.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.