Skip to content
All library documents

Research Protocols for Machine Learning and Quantitative Backtests

Article Hudson & Thames

Summary

This document reviews research practices for applying machine learning and quantitative methods to investing. It outlines common barriers to financial machine learning, including the interdisciplinary nature of the work, limited data, and markets shaped by unpredictable human behavior. It also summarizes a backtesting protocol covering economic rationale, multiple testing, data choices, validation, model changes, complexity, and research culture.

A related discussion of quantitative investing highlights pitfalls involving outliers, normalization, signal decay, turnover, trading costs, rebalancing, short availability, factor payoffs, and portfolio construction. It describes a tutorial using a real-world example, but the document does not provide its detailed results or the full set of pitfalls. These are summaries of cited work rather than new empirical tests. The guidance stresses that out-of-sample performance is difficult to establish before live trading and that researchers should account for failed experiments as well as successful ones.

Key ideas

  • A credible model should have an economic rationale established before testing begins.
  • Researchers should record unsuccessful trials and account for multiple testing.
  • Data selection, transformations, and exclusions need prior justification and robustness checks.
  • Validation should reflect live conditions, including trading costs and data revisions.
  • Simple, interpretable models and a research culture that accepts failed tests can reduce overfitting.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.