Skip to content
All library documents

Generating Trading Ideas and Avoiding Overfit from Data Mining

Article Quant Q&A · Author: confused

Summary

The document asks how trading firms generate research ideas before examining data, and whether they rely heavily on data mining. The question mentions arbitrage, pairs, mean reversion, tick data, momentum, and factor models, while noting practical data limits such as closing prices recorded at different times across markets. It seeks research approaches relevant to professional quant work.

The answer focuses on avoiding in-sample anecdotes from unconstrained models. It recommends regularization, such as penalization, and selecting variables and preprocessing steps to reflect an explicit hypothesis. Constraints can make findings more robust, while a useful explanation should account for why a pattern works and be testable, where possible, with simpler models. Tick-data research examples illustrate a progression from predictive modeling toward explaining price formation. No performance results or general survey of industry idea generation are provided, so these principles do not establish that any particular strategy will work out of sample.

Key ideas

  • Trading research can begin with an economic or market hypothesis that guides variable selection and preprocessing.
  • Regularization methods such as Lasso and Ridge can constrain models and reduce overfitting risk.
  • A predictive pattern is more informative when a model explains the mechanism behind it.
  • Simpler models can help test explanations suggested by more complex tick-data models.
  • The answer addresses overfitting but does not survey how firms generate ideas or prove strategy profitability.

Tags

Full text
# How do you formulate trading ideas and strategies?


# How do you formulate trading ideas and strategies?












I have access to some tick data and Bloomberg data. Outside of data mining and hoping to find an economic rationale after the fact, what do you usually do to generate ideas before you look at the data? I'm having trouble thinking of anything outside of arbitrage, trading pairs of similar assets, data mining correlated assets, data mining stuff that mean reverts, and data mining "signals" from tick data. Maybe I'm just overthinking and data mining actually works. What are some general strategies that firms actually employ as opposed to strategies fit for the personal trader?

I've heard about stuff on momentum (last 11 months skip a month) and the typical "factor" models like small - big, liqudity, but that seems more suited for like mutual funds.

I guess another question is, do firms actually data mine a lot? I'm trying to get a job so I want to do something that is relevant/practical that I can talk about in an interview as opposed to just "I took a stats class or watched a video on ML".

Should I be digging into random research papers? I mean I have ideas in general, but I only have access to some tick data and Bloomberg - which is pretty limited since it only offers closing values which can be a different timestamp for different products/exchanges.

Thanks.

## Answer by lehalle (score 3)

https://quant.stackexchange.com/a/77796

It is a very broad question, let me narrow it and answer on overfitting: how can you prevent blind applications of black box models to exhibit only in sample anecdotes?

First of all, it is worthwhile to focus on two terms

- black box model is not mandatorily a deep neural network, it can be a sequence of moving averages, ranks, and truncations (that is indeed very non linear); a model is a black box when it cannot be explained? here we meet the ambition of Explainable AI. This is not that new: when you use mutual information the measure the relation between two variables (one you observe, one you want to predict), you can end up with the answer that "yes, there is a relation", without being able to exhibit an operative model....

- anecdotes are simply descriptive and not explanatory. It is like witnessing that "you observed a new star in the sky", without giving an understanding of the trajectories of all stars of this kind (to take an image that will speak to econophysicists). This has been theorised by Thomas Kuhn in The Structure Of Scientific Revolutions, be has been also formulated by Shannon and Kolmogorov a more formal way: knowledge has to do with "compression of information" whereas anecdotes are repetition of information.

How van you prevent these issues?



- Regularize your model either by standard penalisation statistical learning uses for long (see Lasso or Ridge), either by a story that explains what you think you leverage on (the story should be reflected in the variables you select, the way you preprocess and mix them). The more constraints you put around your modelling, the more robust will be anything you find despite the constraints.

Let me finish by an example (on tick data), the two papers:

- Sirignano, Justin A. "Deep learning for limit order books" Quantitative Finance 19, no. 4 (2019): 549-570.

- Sirignano, Justin, and Rama Cont. "Universal features of price formation in financial markets: perspectives from deep learning" In Machine Learning and AI in Finance, pp. 5-15. Routledge, 2021.

The first one is a very well done black-box model to predict some features of order book dynamics. And the second (Justin being (co)-author of both, that is why the example is interesting), explain why it works. You can consider the first as a technical trial, and the second as a finite product. Of course what is the best is when the explanation has been formulated before the explorations, or when the explanation are so string that they can be (roughly) tested with simpler models.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.