Genetic Algorithms in Trading: Overfitting and Robust Validation
Summary
The document surveys competing views on using genetic algorithms (GAs) to find financial trading rules or construct portfolios. The main concern is that searching many combinations can produce rules that fit historical data by chance, with data snooping and limited explanatory value. Contributors describe examples of strategies that looked successful in-sample but failed on future data, alongside views that evolutionary methods can be useful when the problem has meaningful structure or when adaptation, rather than prediction, is the goal.
Suggested safeguards include testing on a separate stock universe and delaying deployment while a strategy undergoes further backtesting. These methods can provide additional evidence, but the discussion does not establish a universal validation standard or consensus on GAs. Whether a GA is appropriate depends on the market assumptions, the structure of the task, implementation choices, and the quality of available data. A portfolio optimization anecdote and philosophical arguments are included, but neither demonstrates that GA-derived strategies will generalize reliably.
Key ideas
- Searching many candidate rules with a GA can create overfitting and data snooping risks.
- Validate candidate strategies on data or securities that were kept separate from the search process.
- Some practitioners describe a waiting period and renewed backtesting before allowing a GA-derived strategy into production.
- The usefulness of a GA depends on the problem structure and assumptions about markets, and the discussion gives no single consensus.
Tags
Full text
# How useful is the genetic algorithm for financial market forecasting? # How useful is the genetic algorithm for financial market forecasting? There is a large body of literature on the "success" of the application of evolutionary algorithms in general, and the genetic algorithm in particular, to the financial markets. However, I feel uncomfortable whenever reading this literature. Genetic algorithms can over-fit the existing data. With so many combinations, it is easy to come up with a few rules that work. It may not be robust and it doesn't have a consistent explanation of why this rule works and those rules don't beyond the mere (circular) argument that "it works because the testing shows it works". What is the current consensus on the application of the genetic algorithm in finance? ## Answer by chrisaycock (score 38, accepted) https://quant.stackexchange.com/a/547 I've worked at a hedge fund that allowed GA-derived strategies. For safety, it required that all models be submitted long before production to make sure that they still worked in the backtests. So there could be a delay of up to several months before a model would be allowed to run. It's also helpful to separate the sample universe; use a random half of the possible stocks for GA analysis and the other half for confirmation backtests. ## Answer by vonjd (score 25) https://quant.stackexchange.com/a/548 I think the biggest problem that genetic algorithms have are overfitting, data snooping bias and that they are black boxes (not so much like Neural Networks but still - it depends on the way they are implemented). I think they are not used very much. I guess there are a few hedge funds out there that use it but all in all they were hyped and then busted. (But they are still useful for getting a paper accepted ;-) BTW: There is never a real consensus in finance - everybody tries to outsmart everybody else. This is why it is so interesting. (Or put another way: this is why there are still buyers AND sellers - a real consensus is a crash ;-) ## Answer by bill_080 (score 18) https://quant.stackexchange.com/a/553 I've applied GA to all sorts of things. I had some success in the deterministic world where a pattern actually existed and I knew that some physical structure existed (seismic analysis, vibration analysis, inventory calcs, etc). After I found a GA model that behaved, the real work started....figuring out why it behaved. I also generated a lot of GA garbage from financial data that "worked" looking backward, but was worthless looking forward. Techniques aren't the issue in finance, it's the structure. And, of course, never enough data (useful data). ## Answer by BioinformaticsGal (score 18) https://quant.stackexchange.com/a/913 There's a lot of people here talking about how GAs are empirical, don't have theoretical foundations, are black-boxes, and the like. I beg to differ! There's a whole branch of economics devoted to looking at markets in terms of evolutionary metaphors: Evolutionary Economics! I highly recommend the Dopfer book, The Evolutionary Foundations of Economics, as an intro. http://www.cambridge.org/gb/knowledge/isbn/item1158033?site_locale=en_GB If your philosophical view is that the market is basically a giant casino, or game, then a GA is simply a black-box and doesn't have any theoretical foundation. However, if your philosophy is that the market is a survival-of-the-fittest ecology, then GA's have plenty of theoretical foundations, and it's perfectly reasonable to discuss things like corporate speciation, market ecologies, portfolio genomes, trading climates, and the like. ## Answer by Joshua Chance (score 10) https://quant.stackexchange.com/a/552 Assuming you avoid data-snooping bias and all the potential pitfalls of using the past to predict the future, trusting genetic algorithms to find the "right" solution pretty much boils down to the same bet you make when you actively manage a portfolio, whether quantitatively or discretionary. If you believe in market efficiency then increasing your transaction costs from active management is illogical. If, however you believe there are structural & psychological patterns or "flaws" to be exploited and the payoff is worth the time and money for researching and implementing a strategy the logical choice is active management. Running a GA derived strategy is an implicit bet against market efficiency. You're basically saying "I think there are mis-valuations that occur from some reason" (masses of irrational people, mutual funds herding because of mis-aligned incentives, etc.) and "running this GA can sort this mass of data out way quicker than I can." ## Answer by Greg Thatcher (score 7) https://quant.stackexchange.com/a/18238 I just made a Genetic Algorithms calculator you can try at http://www.gregthatcher.com/Stocks/GeneticAlgorithmCalculator.aspx I'm not a "quant expert" like all of you (I'm just a programmer), but here is what I've found. 1.) If you set the constraints up correctly, the results are amazing. e.g. you can get portfolios that have very high return and low risk. However, it is very important to have conflicting constraints (e.g. a parent can have many children, but the total number of children in a generation cannot go over a certain number) if you want to get good results. 2.) I don't think GA is over-fitting data. Rather, it says "I have too many genes (stocks) to start with, so I'm just going to pick a few to start with, and, except for an occasional mutation, I'll stick with these." Then, over generations, it figures out how to make the best use of what it started with, creating optimum porfolios with the "genes" (a.k.a) stocks it started with (plus a few mutations). Kind of like a builder at Home Depot. Home Depot has lots of tools, but the builder only picks a few to start. IMHO, Genetic Algorithms are an incredible tool for solving problems that human brains can't. ## Answer by RockScience (score 3) https://quant.stackexchange.com/a/546 if you backtest properly your GA (using only past data to generate the time serie of indicator), then you can trust the result. But I agree with you that genetic algorithms are purely empirical and thus I don't feel very comfortable using them. ## Answer by Jagra (score 2) https://quant.stackexchange.com/a/8470 The late Thomas Cover , (likely the leading "Information Theorist" of his generation), considered "Universal" approaches to things like data compression and portfolio allocations as true genetic algorithms. Evolution has no parameters to fit or train. Why should true genetic algorithms? Universal approaches make no assumptions about the underlying distribution of data. They make no attempt to predict the future from patterns or anything else. The "theoretical" effectiveness of Universal approaches (they present significant implementation challenges see my recent question: Geometry for Universal Portfolios?) follow from them doing what evolution demands. The fastest, smartest, or strongest don't necessarily survive in the next generation. Evolution favors that gene, organism, meme, portfolio, or data compression algorithm positioned to most easily adapt to whatever happens next. Also, because these approaches make make no assumptions and operate non-parametrically, one can consider all tests, even on all historical data, as out-of-sample. Certainly they have limitations, Certainly they can't work for every kind a problem we face in our domain, but gee, what an interesting way to think about the things. ## Answer by user533 (score 0) https://quant.stackexchange.com/a/645 Well, the goal of a genetic algo is to find the best solution without going through all the possible scenarios because it would be too long. So of course it is curve fitting, that's the goal!!!
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.