Choosing Factors, Models, and Random Seeds for Equity Prediction
Summary
This practitioner note argues that both feature selection and model choice shape quantitative strategy performance. It compares Random Forest, scikit-learn gradient boosting and random forests, and XGBoost, describing tradeoffs in robustness, speed, nonlinear fit, missing-value handling, tuning burden, and overfitting risk. In the author’s particular exercise, many scikit-learn models had higher annualized returns than XGBoost, attributed to limited sample size and feature dimensions; this is presented as context-specific rather than a general ranking.
The note also gives a reproducibility checklist for randomness across data sampling, feature processing, model training, validation, and evaluation. It recommends fixing seeds for randomized sampling, stochastic estimators, and hyperparameter searches, while preserving ordered splits for time series. Monte Carlo robustness checks should vary seeds deliberately. The performance claims are not accompanied by detailed methodology or validation evidence, and the checklist’s broad statements may depend on estimator settings and implementation. Financial regime changes and noisy signals remain central limitations.
Key ideas
- Factor choice and model choice can both materially affect backtest outcomes.
- Random Forest, scikit-learn boosting methods, and XGBoost have different tradeoffs in fit, speed, tuning, and overfitting risk.
- The reported advantage of scikit-learn models over XGBoost applies only to the exercise’s data and constraints.
- Fix random seeds wherever sampling or stochastic model steps affect reproducibility.
- Use time-ordered validation for financial series and vary seeds intentionally in Monte Carlo robustness checks.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.