Handling Stochastic Variation in Machine Learning Stock Picks
Summary
The discussion addresses how repeated runs of stochastic optimization can produce changing stock rankings, even after the model has converged. Suggested approaches include running multiple instances with different random seeds and aggregating their predictions, such as by taking a median. The spread across runs can serve as an indication of prediction uncertainty.
The answers also recommend examining distributions of candidate portfolios and objective values rather than treating one model run as definitive. Different runs may settle at distinct local optima, and equal objective scores can indicate that the selection criterion does not distinguish among solutions; adding constraints may help. Ensembling is another proposal: combine runs of the same algorithm or models from different algorithms, with attention to their validation performance and prediction correlation. These suggestions are not accompanied by empirical results for the questioner’s strategy, so their usefulness depends on validation design and the stability of out-of-sample performance.
Key ideas
- Repeated runs of stochastic optimization can produce different stock selections despite apparent convergence.
- Aggregating predictions across seeds can provide a more stable signal and a measure of dispersion.
- Compare objective values and portfolio distributions across runs to diagnose multiple optima and indistinguishable solutions.
- Ensembling models with similar validation performance and less correlated predictions is one proposed way to stabilize predictions.
Tags
Full text
# Dealing with stochastic results of Machine Learning Models # Dealing with stochastic results of Machine Learning Models I'm building stock selection models, and pick top 5 and bottom 5 stocks. Given the variability in Stochastic gradient decent results, they keep changing. One way to get consistent results is to use the random seed, but I'm looking for if there a better way to deal with this. Also how would you interpret the results, i.e. One set of top 5 versus another set of Top 5 picks (3-4 of them are the same, but may differ in ranking). I'm running enough iterations to know this isn't an issue about convergence. ## Answer by SachaTheBrave (score 1) https://quant.stackexchange.com/a/54406 Have you tried to choose an arbitrary number of model, let say 20, each one having its own seed? Then you run your twenty models and use the median of your 20 results as signal. One advantage of that method is that you can also get a confidence estimate of your prediction thanks to the standard deviation of your 20 results. ## Answer by Enrico Schumann (score 1) https://quant.stackexchange.com/a/54555 Stochastic solutions are an unavoidable property of stochastic methods, in particular optimisation methods. See for instance section 3 in A Review of Heuristic Optimization Methods in Econometrics. In general, you cannot get rid of randomness; you need to analyse it, by looking at and analysing distributions (e.g. of portfolios) instead of single numbers. See for instance An Empirical Analysis of Alternative Portfolio Selection Criteria (of which I am a coauthor). Convergence means (at best) that the algorithm has stopped in a local optimum. If you have multiple optima, the algorithm may stop at different optima. Have you compared the objective-functions values of repeated runs? Even if they are the same: it means that the algorithm, or more specifically, your selection criterion (=objective function) cannot differentiate between different solutions. Could you modify the model, e.g. add more constraints? ## Answer by Dhruv Mahajan (score 1) https://quant.stackexchange.com/a/54556 The best bet for you is to use Ensemble Learning, as someone experienced with Kaggle competitions, the best way to replicate good performance on Private Learderboard is to ensemble as many algorithms together. This includes intra and inter ensembling. Intra meaning ensembling same algorithms (e.g Xgboost) but with different tuning parameters. You can chose top 10 parameters by cross-validation results to intra-ensemble. Some participants also intra-ensemble different random seeds of same parameters, taking total number of models to more than 500! Second is inter-ensembling, in this you would ensemble different algorithms (e.g neural net, random forest and xgboost), the way you choose these algorithms is by looking at two things : 1) The cross-validation accuracy for each algorithm should be nearby, 2) The correlation in cross-validation predictions should not be more than 80-90%
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.