Skip to content
All library documents

Bagging, Random Forests, and Boosting for Financial Prediction

Article QuantStart

Summary

The article introduces bootstrap resampling and three decision tree ensemble methods. Bagging fits trees to separate samples drawn with replacement and averages their predictions, aiming to reduce the high variance of individual trees. Random forests add feature subsampling at each split to make trees less correlated. Boosting instead grows small models sequentially, fitting each new tree to residual errors and adding its scaled prediction to the ensemble.

It outlines the methods’ tradeoffs: bagging and random forests can be parallelized and do not overfit merely by adding more trees, while boosting is sequential and requires choices such as tree depth, ensemble size, and learning rate. The practical example compares bagging, random forests, and AdaBoost for predicting Amazon daily returns from three lagged returns, using a train-test split and mean squared error. The excerpt provides no reported comparative results. It also cautions that financial observations are serially correlated, so resampled examples may not be independent and ordinary evaluation can overstate statistical validity.

Key ideas

  • Bootstrap resampling creates multiple training samples from one observed dataset.
  • Averaging bootstrapped decision trees can reduce prediction variance.
  • Random forests also sample features at each split to lower correlation among trees.
  • Boosting fits models sequentially to residual errors and combines scaled predictions.
  • Financial time series dependence can undermine assumptions behind resampling and evaluation.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.