Skip to content
All library documents

Random Forest Classification of Long-Horizon Stock Direction

Article BigQuant

Summary

The article describes a supervised learning approach that classifies whether a stock’s closing price will be higher or lower after a chosen horizon. It smooths historical prices with exponential weighting, then uses six technical indicators—including RSI, stochastic measures, MACD, rate of change, and on-balance volume—as features. A random forest combines independently trained decision trees through majority voting, with out-of-bag observations used to estimate classification error.

The reported experiments cover Apple, Samsung, and GE, with predictions at 30-, 60-, and 90-day horizons. The article reports accuracy in the 85%–95% range for those examples, and says accuracy improves and then stabilizes as more trees are added; longer horizons also had lower reported error. Comparisons on a separate 3M dataset are described as favoring random forests over several other classifiers. These are historical classification results, not proof of tradable returns; the summary does not establish out-of-sample robustness, transaction-cost effects, or protection from data leakage.

Key ideas

  • The target is the sign of the difference between the current close and the close at a later horizon.
  • Six technical indicators are used as inputs to a random forest classifier.
  • Bootstrap sampling and majority voting aggregate predictions from multiple decision trees.
  • Out-of-bag error is presented as an estimate of classification mistakes.
  • Reported accuracy rose toward a plateau as tree count increased, while longer forecast horizons had lower reported error.
  • Classification accuracy alone does not establish profitable performance after costs or robust out-of-sample results.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.